← All posts Data engineering

Monitoring and alerting for data pipelines in production

11 Aug 2026 · 5 min read · Twinslytics

Your ROAS dashboard says one thing on Monday and something completely different on Tuesday. Not because ad performance changed, but because a pipeline silently dropped 40% of your Meta conversion events overnight. Nobody noticed until finance asked why blended CAC looked impossible. This is the normal state of data pipelines without real monitoring, and it's costing DTC brands decisions made on garbage numbers.

Most teams treat pipeline monitoring as an afterthought, something to bolt on after the "real" work of building the pipeline is done. That's backwards. In production, the pipeline isn't done until it can tell you when it's broken.

Why silent failures are the default

Data pipelines fail quietly by design. A source API changes a field name, and instead of erroring out, it just returns null. A rate limit kicks in, and instead of failing the job, it returns partial data. Shopify webhooks drop during high traffic and nothing retries them. None of these trigger a red alert. They just produce numbers that look plausible enough to pass a quick glance.

This matters more in ecommerce than most industries because the stakes are attribution and spend decisions. If your warehouse undercounts orders by 8% for three days, your true ROAS calculation looks better than reality. Someone increases ad spend based on that number. The pipeline didn't crash. It just lied, and nobody was watching for lies.

Monitor data, not just jobs

Most teams start with infrastructure monitoring: did the job run, did it finish, did it throw an exception. That's necessary but insufficient. A job can succeed and still produce wrong data.

You need three layers watching your pipelines:

The third layer is the one almost everyone skips, and it's the one that actually protects your attribution model. A row count can look normal while the revenue field itself is wrong because of a currency conversion bug or a duplicate join. Reconciling against the source of truth — Shopify, Klaviyo, the ad platforms themselves — is the only way to catch that.

Set thresholds that mean something

Alert fatigue kills monitoring programs faster than any technical failure. If your Slack channel gets a warning every time row counts shift by 2%, people mute it within a week, and then the pipeline that actually breaks gets ignored along with the noise.

Thresholds need to reflect actual business volatility, not arbitrary round numbers. If your daily order volume naturally swings 15% between a Tuesday and a Friday, don't alert on 10% swings. Instead:

The goal isn't zero alerts. It's alerts that people trust enough to act on immediately, because they know a false-positive rate close to zero.

Build alerts that point to the fix

An alert that says "pipeline failed" at 3am is nearly useless if it takes 45 minutes of log spelunking to find out why. Every alert should carry enough context to start diagnosing without opening five other tools first.

Good alerts include the specific table or job name, the specific check that failed, the actual value versus the expected range, and a link straight to the relevant logs or run history. If a Meta Ads sync fails because of an auth token expiration, the alert should say exactly that, not just "task failed with exit code 1."

This is also where lineage matters. If your orders table breaks, does anyone know that your CAC dashboard, your cohort retention model, and your email flow triggers all depend on it downstream? Without documented lineage, a single broken source table becomes a multi-day forensic exercise instead of a known, contained blast radius. Tools that track dependencies automatically save real hours here, but even a maintained spreadsheet of "this table feeds these five things" beats nothing.

Treat pipelines like production software

The teams that get monitoring right stop treating data pipelines as scripts and start treating them as production software with the same rigor as the checkout flow. That means:

Over time this builds a monitoring layer that actually reflects your specific failure history instead of generic best practices copied from a blog post. Your pipelines break in your specific ways — a particular API's quirks, a particular ETL tool's edge cases, a particular high-traffic sale day that stresses ingestion. The checks that matter are the ones tuned to those actual failure modes.

The real payoff

Monitoring doesn't make your pipelines more impressive. It makes them boring, which is the entire point. Boring pipelines mean your marketing team trusts the ROAS number without a footnote, your finance team doesn't need to reconcile revenue by hand every month-end, and nobody discovers a broken attribution join three weeks after it started skewing budget decisions. Build the alerts before you need them, tune them until people actually trust them, and the payoff shows up as decisions made on real numbers instead of confident guesses.

Further reading

154 dbt models · 4 brands — Multi-brand data platform

Want a warehouse that survives schema drift?

We build daily pipelines that alert on failure down to the file and row, not reporting that quietly breaks.