Build an AI agent that catches refund-skewed ROAS
Facebook says your campaign returned 4.2x ROAS. Three weeks later, half the orders from that campaign get refunded because the size chart was wrong. Nobody updates the dashboard. The media buyer scales the campaign based on a number that was never real to begin with.
This is the gap between reported ROAS and true ROAS, and it's where a lot of ad budget quietly leaks out. Building an AI agent to catch this isn't about replacing your analyst — it's about catching the skew before it turns into a scaling decision you regret.
Why platform roas lies to you
Ad platforms count a conversion the moment the order is placed. They don't know what happens after checkout. A refund three days, three weeks, or three months later never gets subtracted from the number in Ads Manager. The campaign keeps looking profitable long after the actual revenue disappeared.
This isn't a rounding error. Some categories run refund rates north of 20-30% — apparel with sizing issues, high-ticket items with buyer's remorse, subscription boxes with high first-month churn. If your best-performing campaign by platform ROAS is also pushing your highest-return SKUs, you're funding growth with revenue that's going to reverse itself.
The fix isn't a manual monthly reconciliation in a spreadsheet. By the time someone builds that report, the budget's already been reallocated based on bad numbers. You need something watching continuously and flagging the divergence as it happens.
What the agent needs to see
An agent that flags return-skewed ROAS needs three data streams talking to each other, not sitting in separate tools:
- Order and refund data from Shopify or your OMS — every order, every refund event, timestamped, tied to SKU and original order ID.
- Ad spend and attributed revenue from each platform — Meta, Google, TikTok — pulled at the campaign and ad-set level, not just account totals.
- A join key that survives the gap — usually order ID or a UTM-to-order mapping — so a refund three weeks later can be traced back to the campaign that originally drove it.
Without that third piece, you can see refunds happening and you can see ROAS happening, but you can't connect a specific refund to a specific campaign. That's the part most teams skip, and it's the part that actually matters.
This data needs to live in a warehouse, not stay locked in Shopify's admin or Meta's reporting UI. A simple pipeline — Shopify orders and refunds landing in BigQuery or Snowflake, ad platform spend pulled via API on the same schedule — gives the agent one place to query instead of stitching together exports.
How the agent actually works
Strip away the buzzword and this is a scheduled job with logic layered on top. Here's the shape of it:
- Pull raw numbers daily. Spend and attributed revenue by campaign, plus every refund processed against orders from that campaign in the trailing 90 days.
- Calculate net revenue. Attributed revenue minus refunds equals what actually stuck. Divide by spend to get true ROAS, separate from platform-reported ROAS.
- Compare the delta. If platform ROAS says 4x and true ROAS says 2.6x, that's a 35% gap. Set a threshold — say, anything over 20% divergence — that triggers a flag.
- Trace it to a cause. This is where the agent earns its keep. It doesn't just say "ROAS is off," it queries which SKUs or product categories inside that campaign are driving the refunds, and whether it's a handful of orders or a pattern.
- Push the alert somewhere people look. Slack channel, not another dashboard nobody opens. "Campaign X: true ROAS is 2.6x vs reported 4.0x, driven by 60% return rate on SKU-1234" is a message someone can act on same-day.
None of this requires a large language model to function — it's SQL and thresholds. Where an LLM adds value is in the summarization and reasoning layer on top: turning ten flagged campaigns into a two-sentence explanation of what's happening and what to check next, or answering follow-up questions like "has this SKU always had a high return rate, or did something change?"
Rules the agent should follow
A few things separate a useful agent from a noisy one that gets muted after a week:
- Account for return windows by category. A 30-day return window on apparel means you can't finalize true ROAS for a campaign until 30+ days have passed. Flag campaigns as "provisional" until the window closes, then recalculate.
- Separate refunds from returns. A refund without a physical return (customer service goodwill, price adjustment) behaves differently than a full product return. Tag them separately so the agent isn't lumping unrelated issues together.
- Watch for repeat offenders at the SKU level, not just campaign level. Sometimes the campaign is fine — one SKU inside it is the problem. The agent should be able to say "this specific product is dragging every campaign it appears in," which points marketing toward a product fix instead of a media buying fix.
- Don't flag on volume alone. A campaign with $200 in spend and one refund isn't a signal — set a minimum spend or order-count threshold before the agent bothers alerting.
- Recompute, don't just alert once. True ROAS should be a living number that updates as new refunds come in, not a one-time snapshot from the day someone ran the report.
What this changes downstream
Once this agent is running, the immediate value isn't the alert itself — it's what stops happening because of it. Media buyers stop scaling campaigns that look good on day one and fall apart on day thirty. Merchandising gets a clean signal on which SKUs are quietly costing more in returns than they earn in ad-driven sales. Finance gets a ROAS number that matches what actually landed in the bank account, which matters a lot more when it's time to justify budget.
The bigger shift is cultural: ROAS stops being a single number pulled from one platform and becomes something calculated from your own data, checked against reality, and treated with appropriate suspicion until the return window closes. That's a harder habit to build than the pipeline itself.
Start small. Even a weekly SQL job that compares attributed revenue to net revenue after refunds, with output dropped in Slack, beats not knowing at all. Add the automation, the thresholds, and the SKU-level tracing once you've confirmed the gap is real and worth building for. The goal isn't a fancier agent — it's a ROAS number you can actually trust when you decide where the next dollar goes.