A two-hour outage leaves a million events in the outbox. Flushing them is how you turn one incident into two. Why a fixed drain rate is a guess, and how to build a closed-loop drain in Spring Boot that lets downstream latency set the pace.
Counting errors is the wrong signal — it alerts on your busiest merchant every afternoon and stays silent when a bank route quietly dies. A windowed failure rate in Kafka Streams and Spring Boot, and the operational details that decide whether it survives production.
The dual-write bug survives @Transactional, and it quietly loses money. Here is the architecture that fixes it at a million transactions a second — a hash-partitioned outbox on Oracle, SKIP LOCKED relays and Kafka, in Java 21 and Spring Boot — plus some honest scrutiny of the numbers.