
For three years our ingestion pipeline was a series of good decisions stacked on top of each other. Each one made sense at the time. Together they made a system that nobody could reason about end to end.
The moment it stopped scaling
The breaking point was not traffic. It was a Tuesday afternoon when a customer asked why a record from 09:14 appeared in their dashboard at 09:51. Four engineers spent two days answering that question. That is when we knew the architecture had become the product risk.
We ran a two-week audit before writing any code. The finding was uncomfortable: seventy percent of our latency came from a retry layer that existed to work around a bug we had fixed eighteen months earlier.
What we tore out
We removed the retry layer, collapsed three queues into one, and moved schema validation from the consumer to the producer. The rewrite touched eleven services and took nine weeks, which was three weeks longer than we estimated and about half what we feared.
The decision we would make again
We kept the old pipeline running in parallel for six weeks, comparing outputs record by record. It cost us real money in duplicated compute. It also caught two correctness bugs that no test suite would have found, and it meant the cutover was a configuration change rather than an event.