A two-hour outage leaves a million events in the outbox. Flushing them is how you turn one incident into two. Why a fixed drain rate is a guess, and how to build a closed-loop drain in Spring Boot that lets downstream latency set the pace.
Lag is climbing, so you raise the concurrency. Now it is worse. Why consumer count is the wrong lever when downstream is the constraint, how slow processing triggers a rebalance death spiral, and what to tune instead in Spring Boot.
Counting errors is the wrong signal — it alerts on your busiest merchant every afternoon and stays silent when a bank route quietly dies. A windowed failure rate in Kafka Streams and Spring Boot, and the operational details that decide whether it survives production.
Running cache, queue, locks, sessions and rate limiting on a single Redis is the right call for most young systems. It also quietly builds one failure domain spanning all of them — and one of those jobs fails as silent data corruption rather than an outage.
The dual-write bug survives @Transactional, and it quietly loses money. Here is the architecture that fixes it at a million transactions a second — a hash-partitioned outbox on Oracle, SKIP LOCKED relays and Kafka, in Java 21 and Spring Boot — plus some honest scrutiny of the numbers.
Once your data spans services, ACID stops at the service boundary. Sagas trade atomicity for availability — but the compensations, the outbox, and the isolation anomalies are where the real engineering is.