← All work Case study · Data engineering

Cutting a nightly scan by 95% and ending the timeouts

Case study · Data engineering · Twinslytics

A multi-brand data warehouse — 5 brands, one dbt project, 230+ models — kept failing on a nightly timeout. Two months of raising the timeout hadn't fixed it, because the real cause wasn't slow models. It was every model scanning the entire history, every single night.

The problem

The nightly dbt run kept failing on a 3,600-second timeout — but a different model every night, green again after a retry, so for two months the fix was "raise the timeout." The real cause was structural: none of the source tables were partitioned, so every model rescanned the entire history every night, and a handful of heavy models fought each other for BigQuery slots. Other jobs running in the same project made the picture look random.

What we did

The result

Nightly scan on the brand models dropped from 584.6 GiB to 27.0 GiB — a 95.4% reduction — and the full nightly run went from 1.63 TiB to about 1.09 TiB. Nightly timeouts stopped happening entirely, and BigQuery cost fell in proportion to the scan. The output didn't change: equivalence was proven by comparing values, not just row counts.

584.6 → 27.0 GiB nightly scan (−95.4%) · full run 1.63 → ~1.09 TiB · 0 timeouts

Nightly pipeline timing out for a reason nobody's found yet?

We measure the actual scan before touching anything, then fix the structural cause — not the symptom.