9 min readArchitecture
Your Cache Is Load-Bearing
Redis becomes unreachable for forty seconds. Not down — just a failover, or a network partition, or a node restart. Everything is designed for this: Redis is not the source of truth, so the application falls back to the database.
Ninety seconds later the database is dead, and it stays dead well after Redis has come back.
The interesting part is that nothing in the fallback logic was wrong. if (cached == null) return repository.find(id); does exactly what it says. The problem is what happens when every request executes that line at the same moment.
Nobody decided the cache was load-bearing
Start with the number that matters: your cache hit rate. Say it is 90%.
That means the database currently serves 10% of read traffic. Whatever capacity it has — connection pool size, CPU, IOPS — has been sized, tuned and normalised against that 10%. Over months, as the hit rate improved and someone cached one more query, the database’s share shrank and nobody re-provisioned it upward.
So when Redis goes away, read traffic to the database does not increase by 10%. It increases by a factor of ten. A system comfortably handling 2,000 reads per second is suddenly asked for 20,000.
This is the moment the cache stops being a performance optimisation and reveals itself as a load-bearing dependency. Nobody made that decision. It arrived gradually, one @Cacheable at a time, and the architecture diagram still shows Redis as an accelerator sitting politely off to one side.
The database does not degrade gently under that. It hits max_connections, refuses new connections, and every service sharing it starts failing — including the ones that had nothing to do with the cached read. I wrote about that failure mode in detail in the connection exhaustion post; the relevant part here is that it is not gradual. You are fine, and then everything is down at once.
The stampede makes it worse than 10×
Ten times the load would be survivable with headroom. What actually arrives is worse, because of what those requests are.
Consider a popular key — a product, an exchange rate, a merchant configuration — requested 10,000 times a second and served from cache. Redis drops. Now 10,000 concurrent requests independently discover a miss, and each one independently decides to query the database.
They are all asking the same question. The correct number of database queries is one. You are issuing ten thousand.
This is the cache stampede, and it is the difference between a fallback that works and one that amplifies. It also explains the shape people find confusing afterwards: the database was not merely busy, it was executing thousands of copies of the same query, each holding a connection, each waiting behind the others.
Fix one: only one thread may miss
The highest-leverage change is to collapse those 10,000 identical loads into one. Spring’s cache abstraction has this built in and it is a single attribute:
@Cacheable(cacheNames = "merchantConfig", sync = true)
public MerchantConfig findById(String merchantId) {
return repository.findById(merchantId).orElseThrow();
}
With sync = true, concurrent invocations for the same key block while one thread computes the value; the rest receive its result. Ten thousand misses become one query.
Two caveats that matter more than the feature itself.
It is per-JVM, not cluster-wide. The lock lives inside that application’s CacheManager. With 40 pods you get 40 concurrent queries, not one — but 40 is a number a database can absorb, and 10,000 is not. That is usually the whole win, and it is worth being precise about it rather than believing you have global coordination that you do not.
Support depends on the cache provider. Spring’s core CacheManager implementations support it; whether your specific provider does is worth verifying rather than assuming. If it does not, coalesce explicitly:
private final ConcurrentMap<String, CompletableFuture<MerchantConfig>> inFlight =
new ConcurrentHashMap<>();
public MerchantConfig findById(String merchantId) {
MerchantConfig cached = cache.get(merchantId);
if (cached != null) return cached;
// computeIfAbsent is atomic: exactly one loader per key, everyone else
// joins the same future.
CompletableFuture<MerchantConfig> future = inFlight.computeIfAbsent(
merchantId,
key -> CompletableFuture
.supplyAsync(() -> loadAndCache(key), loaderPool)
.whenComplete((v, e) -> inFlight.remove(key)));
return future.join();
}
The whenComplete removal is not optional — without it the map grows forever and you have built a memory leak into your outage-protection code.
Fix two: serve stale rather than nothing
Coalescing reduces the queries. It does not remove them, and during a Redis outage the misses keep coming for every key.
The stronger move is to keep serving data you already have. Split the TTL in two: a soft expiry after which the value is considered stale but still usable, and a hard expiry after which it is genuinely gone.
public record Cached<T>(T value, Instant refreshedAt) {
boolean isStale(Duration softTtl) {
return Instant.now().isAfter(refreshedAt.plus(softTtl));
}
}
public MerchantConfig get(String merchantId) {
Cached<MerchantConfig> entry = cache.get(merchantId);
if (entry == null) {
return loadCoalesced(merchantId); // nothing to serve; must load
}
if (entry.isStale(SOFT_TTL)) {
// Return immediately, refresh in the background, one loader per key.
refreshAsync(merchantId);
}
return entry.value();
}
Now expiry stops being a cliff. When the refresh path is slow or unavailable, users continue to be served — slightly old data instead of an error page. The database sees a trickle of background refreshes rather than a wall of synchronous demand.
This also fixes a self-inflicted problem that has nothing to do with Redis failing. If you warm the cache with a fixed TTL, thousands of keys written at the same moment expire at the same moment, and you stampede yourself on a perfectly healthy Tuesday. Jitter the TTL — a random 10-20% spread — and that disappears:
Duration ttl = BASE_TTL.plusSeconds(
ThreadLocalRandom.current().nextLong(0, BASE_TTL.toSeconds() / 5));
Fix three: the fallback needs a ceiling
Coalescing and stale-serving reduce the load. Something still has to enforce a limit, because the database’s capacity is finite regardless of how well-behaved the callers are.
The instinct is to queue: hold requests until a connection frees up. That is precisely wrong. A queue converts an overload into unbounded latency, and callers respond to slow requests by timing out and retrying — which adds load to the thing that was already saturated.
Bound the fallback explicitly, and reject beyond it:
// Sized to what the database can actually take from THIS service, not to
// how many requests might arrive.
private final Semaphore fallbackPermits = new Semaphore(20);
public MerchantConfig loadFromDatabase(String merchantId) {
// Fail fast. Never block indefinitely waiting for a permit.
if (!fallbackPermits.tryAcquire()) {
throw new FallbackSaturatedException(merchantId);
}
try {
return repository.findById(merchantId).orElseThrow();
} finally {
fallbackPermits.release();
}
}
tryAcquire() without a timeout is the important detail. It fails immediately rather than queueing. The database gets at most 20 concurrent reads from this service and stays responsive; excess requests fail quickly and predictably instead of all of them being slow.
That is the trade this whole design makes: some requests fail cleanly so the rest succeed. Left unbounded, all of them fail together and take the database with them.
The decision nobody writes down
Every pattern list stops at the mechanisms. The judgement call is this, and it is per-read, not per-system:
Is stale data worse than no data?
For a product catalogue, a merchant configuration, an exchange rate a few minutes old — stale is obviously better. Serve it and refresh quietly.
For an account balance, a fraud rule, an authorisation limit — stale may be actively dangerous. Serving a five-minute-old balance during an incident can let money move that should not. Here failing is the correct behaviour, and the honest answer is an error, not a comfortable-looking wrong number.
So do not apply one caching policy across the application. Classify reads by what a stale answer costs, and let that decide the soft TTL, whether stale-while-revalidate applies at all, and whether the fallback is even permitted to be shed. Two lines in a config, but it is a business decision and it needs someone from the business in the room.
When Redis is down, stop asking it
One more failure mode, distinct from a miss. If Redis is unreachable rather than empty, every single request pays the connection timeout before falling back. A 1-second timeout at 10,000 requests per second does not just add latency — it exhausts your thread pools while achieving nothing.
Keep the timeouts brutally short and put a breaker in front:
spring:
data:
redis:
timeout: 150ms # a healthy Redis answers in single-digit ms
connect-timeout: 200ms
@CircuitBreaker(name = "redis", fallbackMethod = "loadWithoutCache")
public MerchantConfig get(String merchantId) { ... }
Once the breaker opens, requests skip Redis entirely and go straight to the (now bounded, coalesced) fallback. You have removed a dead dependency from the hot path instead of paying for it on every request.
Also cache your misses. If a key genuinely does not exist and you only cache positive results, every request for it reaches the database forever — cache penetration, and it is trivially weaponisable if the key comes from user input. A short-TTL negative entry costs almost nothing and closes it.
Test the thing you are afraid of
All of this is unverifiable by reading it. The only way to know whether your fallback holds is to remove Redis while the system is under realistic load, and watch what the database does.
Do it in staging, with production-like traffic, with everything switched on. The failures you find will be specific and unglamorous: a @Cacheable that someone added without sync, a connection pool sized for the happy path, a timeout still at its 30-second default, a retry policy that triples the load exactly when it should back off.
Every one of those is cheap to fix on a Tuesday and expensive to discover during an incident.
The short version
- Work out your hit rate. That number tells you the multiple your database will absorb when the cache disappears.
- A cache above ~80% hit rate is a dependency, not an optimisation. Treat it in your capacity planning and your architecture diagram accordingly.
- Coalesce misses —
@Cacheable(sync = true), remembering it is per-JVM. - Serve stale while revalidating, so expiry is not a cliff.
- Jitter your TTLs, or you will stampede yourself without any help from Redis.
- Bound the fallback and fail fast. Never queue into a saturated database.
- Decide per read whether stale beats an error. It is a business question.
- Break the circuit on Redis itself, so a dead cache is skipped rather than waited on.
- Cache negative results to close the penetration path.
- Kill Redis in staging under load. Nothing else tells you the truth.
Cache for speed, database for truth — but the sentence needs a third clause. Protect the path between them, because that is the one that carries all the traffic on the worst day.