Guide 02
Brokers and frameworks: Celery, Redis, RabbitMQ, Kafka
Audience: Architects and leads picking a worker stack. Policy: Durable jobs use RabbitMQ and/or Kafka. Redis is not the primary job broker.
TL;DR
| Piece | Role here |
|---|---|
| RabbitMQ | Commands / jobs (acks, routing, DLX) |
| Kafka | Domain events / streams (fan-out, lag, replay) |
| Redis | Cache, rate limits, optional Celery result backend |
| Celery / Taskiq / Dramatiq | Job frameworks on AMQP |
| FastStream | Event consumers/producers on Kafka (or RMQ) |
| ARQ / RQ | Redis queues : not system of record under this checklist |
Contents
- Roles
- Why Redis is demoted
- Stack matrix
- Topology patterns
- Framework comparison
- Config checklist
- Anti-patterns
---
1. Roles
Keep transport separate from product state:
- Broker moves messages.
- Your jobs (or projections) table is what the product UI reads.
- Flower / RMQ UI / Kafka UI are ops tools, not user-facing status.
2. Why Redis is demoted for durable jobs
Redis lists are fast. They are a weak default when you need:
- Durable queues across broker restarts with mature ops tooling
- Dead-letter exchanges and redrive
- Prefetch / QoS for mixed job sizes
- Clear separation of job traffic from cache traffic
OK uses of Redis: session cache, rate limits, feature flags, optional Celery result backend.
Not OK as only home of payment email / invoice PDF: business-critical commands belong on RabbitMQ (or an explicit Kafka design with care).
3. Stack matrix
| Need | Pick |
|---|---|
| Classic jobs, Beat, Flower, team knows Celery | Celery + RabbitMQ |
| Async FastAPI + DI-friendly tasks | Taskiq + RabbitMQ |
| Simple reliable sync actors | Dramatiq + RabbitMQ |
| Domain events, fan-out, replay | FastStream + Kafka |
| Jobs + events | RMQ workers + Kafka consumers |
| After-response fluff only | BackgroundTasks |
Decision tree
text
Business-critical or must survive deploys?
NO → BackgroundTasks (tiny)
YES → Is it a command (one worker does work) or an event (many react)?
COMMAND → RabbitMQ + Celery | Taskiq | Dramatiq
EVENT → Kafka + FastStream (often after transactional outbox)
BOTH → RMQ for jobs + Kafka for events4. Topology patterns
Jobs only
text
API → RabbitMQ (jobs.*) → workers → DB / email / S3Events only
text
API/service → outbox → Kafka topic → consumers (own DBs)Combined (recommended for growing products)
text
API
├─ command → RabbitMQ → job worker (export, email)
└─ outbox event → Kafka → notifications, analytics, ledger5. Framework comparison
| Celery | Taskiq | Dramatiq | FastStream | |
|---|---|---|---|---|
| Primary model | Tasks | Async tasks | Actors | Streams |
| RabbitMQ | Excellent | Yes | Yes | Yes |
| Kafka | Limited/community | Yes (broker) | No (use other) | Excellent |
| Beat / schedule | Mature | Via schedules | Periodics | App-level |
| Flower-like UI | Flower | Custom / metrics | Custom | Metrics |
| Best fit | Existing Celery shops | Async FastAPI | Simple RMQ workers | Event pipelines |
Deep dives: 05 Celery · 06 Taskiq · 07 Dramatiq · 08 FastStream
6. Config checklist
- [ ] Broker URL from Pydantic Settings (not hardcoded)
- [ ] TLS in production (
amqps://, Kafka SSL) - [ ] Separate vhost / cluster per environment
- [ ] JSON serializers only (no pickle)
- [ ] Queue/topic names versioned or stable + documented
- [ ] DLQ / error topic defined
- [ ] Worker and API as different deployables, same image version
7. Anti-patterns
- Redis-only queue while the org requires RMQ/Kafka
- Pickle serializers
- Treating the broker UI as the product status API
- One mega-queue for 50ms and 50-minute jobs
- API and worker in one container without independent scale
Next: 03 Job lifecycle