Monitoring and alerting for data pipelines in production
Your ROAS dashboard says one thing on Monday and something completely different on Tuesday. Not because ad performance changed, but because a pipeline silently dropped 40% of your Meta conversion events overnight. Nobody noticed until finance asked why blended CAC looked impossible. This is the normal state of data pipelines without real monitoring, and it's costing DTC brands decisions made on garbage numbers.
Most teams treat pipeline monitoring as an afterthought, something to bolt on after the "real" work of building the pipeline is done. That's backwards. In production, the pipeline isn't done until it can tell you when it's broken.
Why silent failures are the default
Data pipelines fail quietly by design. A source API changes a field name, and instead of erroring out, it just returns null. A rate limit kicks in, and instead of failing the job, it returns partial data. Shopify webhooks drop during high traffic and nothing retries them. None of these trigger a red alert. They just produce numbers that look plausible enough to pass a quick glance.
This matters more in ecommerce than most industries because the stakes are attribution and spend decisions. If your warehouse undercounts orders by 8% for three days, your true ROAS calculation looks better than reality. Someone increases ad spend based on that number. The pipeline didn't crash. It just lied, and nobody was watching for lies.
Monitor data, not just jobs
Most teams start with infrastructure monitoring: did the job run, did it finish, did it throw an exception. That's necessary but insufficient. A job can succeed and still produce wrong data.
You need three layers watching your pipelines:
- Job-level monitoring — did the pipeline run, how long did it take, did it error out. This catches crashes and timeouts.
- Data-level monitoring — row counts, null rates, distinct value counts, freshness of the most recent timestamp. This catches silent data loss.
- Business-level monitoring — does total revenue in the warehouse match Shopify's own reported revenue for the day, does order count match, does spend match what the ad platform reports. This catches everything the first two layers miss.
The third layer is the one almost everyone skips, and it's the one that actually protects your attribution model. A row count can look normal while the revenue field itself is wrong because of a currency conversion bug or a duplicate join. Reconciling against the source of truth — Shopify, Klaviyo, the ad platforms themselves — is the only way to catch that.
Set thresholds that mean something
Alert fatigue kills monitoring programs faster than any technical failure. If your Slack channel gets a warning every time row counts shift by 2%, people mute it within a week, and then the pipeline that actually breaks gets ignored along with the noise.
Thresholds need to reflect actual business volatility, not arbitrary round numbers. If your daily order volume naturally swings 15% between a Tuesday and a Friday, don't alert on 10% swings. Instead:
- Build thresholds off trailing averages and standard deviation, not fixed percentages. A day that's three standard deviations off the 28-day rolling average is worth a page. A day that's 5% below yesterday probably isn't.
- Separate warning-level alerts (Slack message, check when convenient) from critical alerts (page someone, this blocks a report or a decision). Mixing them trains people to ignore both.
- Alert on freshness separately from volume. A table that hasn't updated in six hours is a different problem than a table that updated with 20% fewer rows than expected, and they need different urgency.
The goal isn't zero alerts. It's alerts that people trust enough to act on immediately, because they know a false-positive rate close to zero.
Build alerts that point to the fix
An alert that says "pipeline failed" at 3am is nearly useless if it takes 45 minutes of log spelunking to find out why. Every alert should carry enough context to start diagnosing without opening five other tools first.
Good alerts include the specific table or job name, the specific check that failed, the actual value versus the expected range, and a link straight to the relevant logs or run history. If a Meta Ads sync fails because of an auth token expiration, the alert should say exactly that, not just "task failed with exit code 1."
This is also where lineage matters. If your orders table breaks, does anyone know that your CAC dashboard, your cohort retention model, and your email flow triggers all depend on it downstream? Without documented lineage, a single broken source table becomes a multi-day forensic exercise instead of a known, contained blast radius. Tools that track dependencies automatically save real hours here, but even a maintained spreadsheet of "this table feeds these five things" beats nothing.
Treat pipelines like production software
The teams that get monitoring right stop treating data pipelines as scripts and start treating them as production software with the same rigor as the checkout flow. That means:
- Every pipeline has an owner, not just a creator. Someone is on the hook when it breaks, even if they didn't write the original code.
- Every critical table has documented expectations: expected row count range, expected freshness, expected null rate, written down somewhere before the pipeline goes live, not reverse-engineered after an incident.
- Incidents get a short postmortem. Not a blame exercise, just: what broke, what alert caught it (or didn't), what check gets added so it's caught faster next time.
Over time this builds a monitoring layer that actually reflects your specific failure history instead of generic best practices copied from a blog post. Your pipelines break in your specific ways — a particular API's quirks, a particular ETL tool's edge cases, a particular high-traffic sale day that stresses ingestion. The checks that matter are the ones tuned to those actual failure modes.
The real payoff
Monitoring doesn't make your pipelines more impressive. It makes them boring, which is the entire point. Boring pipelines mean your marketing team trusts the ROAS number without a footnote, your finance team doesn't need to reconcile revenue by hand every month-end, and nobody discovers a broken attribution join three weeks after it started skewing budget decisions. Build the alerts before you need them, tune them until people actually trust them, and the payoff shows up as decisions made on real numbers instead of confident guesses.