Checklist/Docs/Brokers and frameworks: Celery, Redis, RabbitMQ, Kafka

Guide 02

Brokers and frameworks: Celery, Redis, RabbitMQ, Kafka

Audience: Architects and leads picking a worker stack. Policy: Durable jobs use RabbitMQ and/or Kafka. Redis is not the primary job broker.

TL;DR

Production layout: API, workers, Beat, Flower, brokers
ClientsAPIFastAPI × NPostgresjobs / outboxRabbitMQcommandsKafkaevents (opt)Redislocks / RL onlyWorkersCelery × NBeat1 leaderFlowerops UI privatePrometheusscrape metricsSame image · different CMD
PieceRole here
RabbitMQCommands / jobs (acks, routing, DLX)
KafkaDomain events / streams (fan-out, lag, replay)
RedisCache, rate limits, optional Celery result backend
Celery / Taskiq / DramatiqJob frameworks on AMQP
FastStreamEvent consumers/producers on Kafka (or RMQ)
ARQ / RQRedis queues : not system of record under this checklist

Contents

  1. Roles
  2. Why Redis is demoted
  3. Stack matrix
  4. Topology patterns
  5. Framework comparison
  6. Config checklist
  7. Anti-patterns

---

1. Roles

Keep transport separate from product state:

  • Broker moves messages.
  • Your jobs (or projections) table is what the product UI reads.
  • Flower / RMQ UI / Kafka UI are ops tools, not user-facing status.

2. Why Redis is demoted for durable jobs

Redis lists are fast. They are a weak default when you need:

  • Durable queues across broker restarts with mature ops tooling
  • Dead-letter exchanges and redrive
  • Prefetch / QoS for mixed job sizes
  • Clear separation of job traffic from cache traffic

OK uses of Redis: session cache, rate limits, feature flags, optional Celery result backend.

Not OK as only home of payment email / invoice PDF: business-critical commands belong on RabbitMQ (or an explicit Kafka design with care).

3. Stack matrix

NeedPick
Classic jobs, Beat, Flower, team knows CeleryCelery + RabbitMQ
Async FastAPI + DI-friendly tasksTaskiq + RabbitMQ
Simple reliable sync actorsDramatiq + RabbitMQ
Domain events, fan-out, replayFastStream + Kafka
Jobs + eventsRMQ workers + Kafka consumers
After-response fluff onlyBackgroundTasks

Decision tree

text
Business-critical or must survive deploys?
 NO → BackgroundTasks (tiny)
 YES → Is it a command (one worker does work) or an event (many react)?
 COMMAND → RabbitMQ + Celery | Taskiq | Dramatiq
 EVENT → Kafka + FastStream (often after transactional outbox)
 BOTH → RMQ for jobs + Kafka for events

4. Topology patterns

Jobs only

text
API → RabbitMQ (jobs.*) → workers → DB / email / S3

Events only

text
API/service → outbox → Kafka topic → consumers (own DBs)

Combined (recommended for growing products)

text
API
 ├─ command → RabbitMQ → job worker (export, email)
 └─ outbox event → Kafka → notifications, analytics, ledger

5. Framework comparison

CeleryTaskiqDramatiqFastStream
Primary modelTasksAsync tasksActorsStreams
RabbitMQExcellentYesYesYes
KafkaLimited/communityYes (broker)No (use other)Excellent
Beat / scheduleMatureVia schedulesPeriodicsApp-level
Flower-like UIFlowerCustom / metricsCustomMetrics
Best fitExisting Celery shopsAsync FastAPISimple RMQ workersEvent pipelines

Deep dives: 05 Celery · 06 Taskiq · 07 Dramatiq · 08 FastStream

6. Config checklist

  • [ ] Broker URL from Pydantic Settings (not hardcoded)
  • [ ] TLS in production (amqps://, Kafka SSL)
  • [ ] Separate vhost / cluster per environment
  • [ ] JSON serializers only (no pickle)
  • [ ] Queue/topic names versioned or stable + documented
  • [ ] DLQ / error topic defined
  • [ ] Worker and API as different deployables, same image version

7. Anti-patterns

  • Redis-only queue while the org requires RMQ/Kafka
  • Pickle serializers
  • Treating the broker UI as the product status API
  • One mega-queue for 50ms and 50-minute jobs
  • API and worker in one container without independent scale

Next: 03 Job lifecycle