9 min readArchitecture
One Redis, Seven Jobs, One Outage
I have written a fair amount here about outboxes, sagas, hash-partitioned relays, backlog drain controllers and stream processing topologies. So let me say the quiet part clearly:
Most teams do not need any of that yet, and building it early is the more expensive mistake.
Premature enterprise architecture is far more common than under-engineering, and it fails in a particularly bad way — not with an outage, but with six months of delivery spent on machinery for problems the business has not encountered. Kafka, a service mesh, CQRS and event sourcing on a system serving four hundred requests a minute is not preparation. It is a tax on every feature you ship until you get there, and many products never do.
So the argument for starting with a single Redis — cache, pub/sub, Streams as a queue, sessions, rate limiting, distributed locks, leaderboards — is sound. One dependency to run, one thing for the team to learn, near-zero operational cost, and you can build all of it in an afternoon.
I agree with all of that. What usually goes unsaid is what you signed up for, and one part of it is genuinely dangerous.
Seven jobs, one failure domain
Consolidation is not just a cost saving. It is a decision to couple the availability of seven unrelated concerns.
When that Redis has a bad thirty seconds — a failover, a network partition, an eviction storm, someone running KEYS * in production — you do not lose one capability. You lose all of them simultaneously, and they fail in quite different ways:
Cache falls back to the database, which has been quietly provisioned for whatever fraction of reads it currently serves. That is its own failure mode, and it is the one people at least anticipate.
Sessions disappear, so every user is logged out at once. Then all of them log in again, simultaneously, against an auth path sized for a normal login rate.
Rate limiting stops working, and now you must choose in advance between two bad options. Fail open, and your APIs are unprotected exactly when the platform is least healthy. Fail closed, and a Redis blip becomes a total outage of everything behind the limiter. Whichever you choose, choose it deliberately — the default is usually “whatever the library does”, which nobody has read.
Queued work stalls or is lost, depending on durability settings we will get to.
And then the one that is qualitatively different.
The lock is not an availability problem
Distributed locks fail as a correctness problem, not an availability problem, and that distinction is the single most important thing in this post.
If Redis is unavailable and your lock acquisition fails closed, you get an outage — bad, but visible, and it stops. If it fails open, or if two nodes disagree about who holds the lock during a failover, then two workers do the same job at the same time. Two settlement runs. Two disbursements. Two charges.
Nothing pages. Nothing errors. The dashboards stay green. You find out days later during reconciliation, and by then the damage is distributed across thousands of records and has to be unwound by hand.
Redis-based locking is fine for things where a rare double-execution is merely wasteful — refreshing a cache entry, sending a “your report is ready” email, running a cleanup job. It is not a safe foundation for anything where double-execution moves money or changes legal state. There, the correctness has to live in the database that owns the data: a unique constraint, a conditional update, a status transition that can only happen once. The lock becomes an optimisation to avoid wasted work, not the thing standing between you and a duplicate payment.
That is a design rule worth adopting on day one, because it costs nothing then and is very expensive to retrofit.
Redis Streams as a queue: what you actually get
Streams are a real queue, and better than people assume. Consumer groups, per-consumer pending entry lists, XACK for acknowledgement, XAUTOCLAIM to recover messages from a consumer that died mid-work, and history you can re-read with XRANGE independently of any group’s cursor. For a large class of background work that is genuinely enough.
The limit is durability, and it is worth being precise rather than hand-wavy, because this is the line that decides what may and may not live there.
Redis replication is asynchronous by default. The primary does not wait for replicas to confirm before acknowledging your write. So if a primary fails over before replicas have caught up, writes that Redis already told you it accepted can be gone.
The WAIT command improves this — it blocks until a given number of replicas acknowledge — but the Redis documentation is explicit that this does not give strong consistency, and that failover remains best-effort and may still lose data under specific failure conditions. An AOF file with a strict fsync policy raises durability further, at a throughput cost.
So the practical boundary:
- Fine on Redis Streams: send the welcome email, generate the thumbnail, refresh the search index, recalculate a leaderboard. Losing one occasionally is recoverable, often invisible, and always fixable by re-running something.
- Not fine on Redis Streams: anything where the message is the money. A payment instruction that vanishes during a failover is not an inconvenience; it is a customer whose transfer disappeared, and no amount of monitoring recovers a message that was never persisted.
That is also why the transactional outbox lives in the database rather than in the broker. The durability guarantee comes from the database transaction, not from the messaging system — which means the outbox pattern works perfectly well with Redis Streams as the transport. You get the safety without needing Kafka.
That combination — outbox in Postgres, Redis Streams as the relay target — is a genuinely good architecture for a young system handling real money, and it costs you one Redis instead of a Kafka cluster.
The question worth asking early
Not “when do we move to Kafka”, which is unanswerable in the abstract. Instead:
Which of these jobs has to keep working when Redis does not?
It takes an hour with a whiteboard and it changes what you build. Sort the seven jobs into three piles:
- Degrades acceptably — cache, leaderboards, non-critical pub/sub. Redis down means slower or slightly stale. Design the fallback, bound it, move on.
- Must fail closed — rate limiting, quota enforcement. Decide the behaviour explicitly and write it down, because the default is usually wrong for you.
- Must not be wrong — locks guarding money, queues carrying financial instructions. These need a correctness guarantee that does not depend on Redis at all, and they are the first things to move when you outgrow the single instance.
Most teams have far less in pile three than they fear. But the things in it are the things that end up in a post-mortem.
Moving is triggered by a property, not a number
There is no request-per-second figure at which you should adopt Kafka. The trigger is needing a specific property you cannot get where you are:
- Durable retention and replay from an arbitrary point — reprocessing last month with fixed logic, or seeding a new service from history.
- Independent fan-out — several consumer groups reading the same stream at their own pace, added later without touching the producer.
- Retention measured in days or weeks, decoupled from memory.
- Ordering guarantees per key at high volume, with partitioning you control.
RabbitMQ sits in the middle for a different reason: routing. Topic exchanges, per-message TTLs, dead-letter exchanges and priority queues are things Redis Streams does not do and Kafka does awkwardly. If your problem is complex routing of tasks rather than a high-volume event log, RabbitMQ is the better fit, and skipping it to go straight to Kafka is its own kind of over-engineering.
Make the exit cheap
The real cost of an early choice is not the technology. It is how much code assumes it.
A thin interface at the boundary costs nothing on day one and makes the eventual move a contained change rather than a rewrite:
public interface EventPublisher {
void publish(String stream, DomainEvent event);
}
@Component
@Profile("!kafka")
@RequiredArgsConstructor
public class RedisStreamPublisher implements EventPublisher {
private final StringRedisTemplate redis;
@Override
public void publish(String stream, DomainEvent event) {
redis.opsForStream().add(StreamRecords.newRecord()
.in(stream)
.ofMap(Map.of(
"type", event.type(),
"payload", serialize(event),
// Carry this from day one. Consumers need it for idempotency
// whatever the transport turns out to be.
"eventId", event.id())));
}
}
The important part is not the interface — it is the eventId and the discipline that goes with it. Consumers that are idempotent, messages that carry their own identity, and no business logic that depends on transport-specific behaviour. Get those right on Redis and the eventual migration is mechanical. Get them wrong and it is a rewrite regardless of how clean your interface looked.
Do not build an abstraction layer that tries to hide the differences between Redis Streams and Kafka. They have genuinely different semantics and a lowest-common-denominator wrapper gives you the weaknesses of both. One narrow interface at the edge, not a framework.
What I would actually do
Early — one Redis for cache, sessions, rate limiting, background jobs via Streams. Money-critical correctness enforced by database constraints, never by a Redis lock. Outbox in the database from the first payment feature, because retrofitting it after you have lost a transaction is far worse than writing it now.
Growing — separate the Redis instances before scaling any of them. Cache on one, queues on another. Same technology, different blast radii, and it is a config change rather than a migration.
At scale — move to Kafka when you need retention, replay or independent fan-out. By then you will have a concrete reason, which is a much better basis for the decision than a diagram.
The short version
- Premature enterprise architecture is the more common mistake. Starting on one Redis is usually right.
- Consolidation couples availability across every job you put on it.
- Locks fail as correctness, not availability. Never let a Redis lock be the only thing preventing a duplicate payment.
- Redis replication is asynchronous. Acknowledged writes can be lost on failover;
WAIThelps and is still best-effort. - Sort your jobs by what happens when Redis is down. An hour’s work, done early.
- The outbox pattern works with Redis Streams — durability comes from your database, not your broker.
- Move for a property, not a number. Retention, replay, fan-out, routing.
- Split instances before you switch technologies. Cheapest risk reduction available.
Good architecture is not the fewest technologies any more than it is the most. It is knowing precisely what each choice costs you on the day it fails — and having decided, in advance and on purpose, that you can live with it.