At a 90% hit rate your database is provisioned for 10% of your traffic. Nobody decided that — it happened gradually. Which is why 'fall back to the DB' turns a 40-second Redis blip into a database outage, and what to build instead in Spring Boot.
A two-hour outage leaves a million events in the outbox. Flushing them is how you turn one incident into two. Why a fixed drain rate is a guess, and how to build a closed-loop drain in Spring Boot that lets downstream latency set the pace.
Lag is climbing, so you raise the concurrency. Now it is worse. Why consumer count is the wrong lever when downstream is the constraint, how slow processing triggers a rebalance death spiral, and what to tune instead in Spring Boot.
Counting errors is the wrong signal — it alerts on your busiest merchant every afternoon and stays silent when a bank route quietly dies. A windowed failure rate in Kafka Streams and Spring Boot, and the operational details that decide whether it survives production.
Connection exhaustion takes down a database that isn't even busy. Here's why the arithmetic guarantees it at scale, and how connection multiplexing with ProxySQL breaks the link between application pools and database threads.