Build an AI agent that catches ROAS anomalies before overspend
By the time your Monday morning dashboard shows a ROAS crater, you've already spent Friday, Saturday, and Sunday's budget on a campaign that stopped converting. Nobody checks ad spend on a Saturday at 2am. An AI agent will. This is a practical build guide, not a pitch for a SaaS tool you don't need.
Why you find out too late
Most teams catch ROAS problems through a person: a media buyer glancing at Meta Ads Manager, a founder asking "why does this month look weak" in a Slack thread. That's a detection system with a 24-72 hour lag built in. Platforms report blended numbers that hide the collapse of a single ad set inside an otherwise healthy campaign. And attributed ROAS from ad platforms is inflated anyway, so the number you're staring at was already wrong before it started drifting.
The fix isn't "check the dashboard more often." It's removing the human from the detection loop and keeping them for the decision loop. An agent watches spend and revenue continuously, flags deviations against a real baseline, and pushes an alert before the budget is gone — not after finance asks about it in the weekly review.
The data layer the agent needs
An anomaly detector is only as good as the numbers feeding it, and this is where most attempts break. You need a pipeline that pulls, at minimum:
- Platform spend data at the campaign and ad set level, pulled hourly via the Meta, Google, and TikTok APIs — not the daily CSV export someone downloads manually.
- Order and revenue data from your store or order management system, ideally with UTM or click ID matching so you're not relying solely on platform-attributed conversions.
- A blended or true ROAS calculation that nets out discounts, refunds, and COGS where possible, computed in your warehouse, not trusted from the ad platform's dashboard.
- Historical baselines going back at least 90 days per channel and campaign, so the agent has something real to compare against instead of an arbitrary threshold.
This lives in a warehouse — BigQuery or Snowflake, with dbt models turning raw API pulls into a clean daily_channel_performance table isn't allowed as a tag here, so just: a clean, tested table joining spend to revenue by channel and day. If that pipeline doesn't exist yet, build it before you build the agent. An agent reasoning over garbage data will just automate garbage alerts faster.
Define anomaly, not just alert
The lazy version of this is a static rule: "alert if ROAS drops below 1.5." Static thresholds break constantly. A new product launch, a holiday promo, a deliberate top-of-funnel push — all of these tank short-term ROAS on purpose. A rule-based system floods you with false positives until someone mutes it, and then it's dead.
What actually works is comparing current performance to a rolling statistical baseline, adjusted for known variance:
- Rolling median and standard deviation per campaign over a trailing 14-30 day window, so the baseline moves with recent reality instead of a fixed number set six months ago.
- Z-score or MAD-based deviation checks that flag when today's ROAS or CPA is a statistically meaningful outlier, not just "a little lower than yesterday."
- Spend velocity checks that catch a campaign burning through its daily budget in three hours instead of eighteen — often the earliest signal something is broken, before ROAS even has time to move.
- Seasonality and calendar context so the agent knows Black Friday week and a new SKU launch aren't anomalies, they're expected variance.
This is the part people skip because it sounds like statistics homework. It's not optional. Without it, you're building an expensive way to get alerted about noise.
Build the detection loop
The agent itself doesn't need to be exotic. A practical version runs as a scheduled job — hourly is usually the right cadence for paid social, every 4-6 hours is fine for search — and does four things in sequence:
- Pull fresh spend and revenue data from the warehouse tables you already built.
- Score each active campaign or ad set against its rolling baseline using the deviation logic above.
- Classify the anomaly by severity — a mild dip in ROAS gets logged, a spend spike with no matching revenue gets an immediate alert.
- Notify through a channel people actually check — Slack or a text alert, with the specific campaign name, the deviation, and the dollar amount at risk, not a generic "something changed" message.
This can run as a Python script on a cron schedule calling an LLM only for the summarization step — turning "campaign X, z-score -2.4, spend up 40% day-over-day, revenue flat" into a readable Slack message a media buyer can act on in ten seconds. The "AI" part doesn't need to be doing the statistics; language models are bad at math and good at explaining a number clearly. Let a proper statistical function do the detection, and let the agent handle interpretation, prioritization, and communication.
For teams that want to go further, the agent can also take a first action — pausing an ad set automatically when spend velocity crosses a hard ceiling with zero attributed revenue, then notifying the team that it did so. That's the difference between an alert system and an agent: it acts, within guardrails you set, instead of just talking.
Keep it from crying wolf
The fastest way to kill this system is alert fatigue. If it pings the team five times a day for normal variance, everyone mutes the channel within a week and you're back to finding out on Monday. A few things keep it credible:
- Tiered alerting — log minor deviations quietly, reserve pings for anomalies that cross a spend-at-risk threshold you actually care about, like $500 or more of unexplained spend.
- Suppression windows during known events — launches, promos, platform-wide auction volatility during major shopping days — so the agent doesn't flag things you already expect.
- Feedback loop — when someone marks an alert as a false positive, feed that back into the baseline logic so the model tightens over time instead of staying static.
- A weekly accuracy review — track how many alerts were real problems versus noise, and adjust thresholds like you would any other model, because that's what this is.
None of this needs to be perfect on day one. Start with spend velocity and one or two high-spend channels, get the alerting cadence right, then expand coverage once the team trusts what it's seeing.
The goal isn't a dashboard with more charts. It's a system that catches the $2,000 you were about to waste on a broken ad set before Saturday turns into Sunday. Build the data pipeline first, define anomalies statistically instead of with arbitrary thresholds, and let the agent do the tedious watching so your team can do the deciding.