Checklist/Docs/Monitoring Celery with Flower

Guide 18

Monitoring Celery with Flower

Audience: Ops and backend engineers running Celery workers on RabbitMQ. Golden rule: Flower is an ops dashboard. Product job status always comes from your `jobs` table / GET /jobs/{id} : never from Flower.

TL;DR

Observability: Flower, Prometheus, app metrics
Workers-E eventsFlower/metricsPrometheusscrape 15sGrafanaAlertmanagerRabbitMQRMQ exporterdepth · DLQAPIGET /jobsproduct status = DB · Flower = ops only
TopicGuidance
What Flower showsWorkers online, active/reserved/scheduled tasks, basic history, rates
What it does not replaceQueue depth/DLQ alerts, app metrics, user-facing job status
How to runcelery -A ... flower as a separate process
SecurityPrivate network + authentication + TLS (mandatory in prod)
BrokerWorks with RabbitMQ (this checklist); Redis broker not our job default
text
Celery workers ◄── events / inspect ──► Flower (ops UI)
 │
 └── still emit metrics/logs to Prometheus/your stack

Contents

  1. What Flower is for
  2. What Flower is not for
  3. Install and run
  4. Essential configuration
  5. Security (mandatory)
  6. What to watch in the UI
  7. Enable Celery events
  8. Flower + metrics (better together)
  9. Deploy sketch
  10. Troubleshooting
  11. Checklist

---

1. What Flower is for

Flower is a real-time web monitor for Celery:

  • List workers (alive, concurrency, queues)
  • See active, reserved, scheduled tasks
  • Inspect task success/failure history (when events are enabled)
  • Basic rates and task runtime views
  • Optional revoke / shutdown controls (treat as dangerous in prod)

Use it during incidents: “Are workers connected? Is a task stuck active? Did failures spike?”

---

2. What Flower is not for

Not thisUse this instead
Customer “is my export done?”GET /jobs/{id} from application DB
Sole source of queue backlog SLOsRabbitMQ depth / consumer count metrics
Public status pageNever expose Flower publicly without auth
Long-term analytics warehousePrometheus + logs + job table
Replace structured loggingJSON logs with job_id / task_id

See 15 Observability.

---

3. Install and run

bash
# same app env as workers
pip install flower

# PSEUDOCODE : module path to Celery app instance
celery -A app.workers.celery_app.celery_app flower \
 --port=5555 \
 --basic_auth=ops_user:strong_password
python
# PSEUDOCODE : app/workers/celery_app.py must be importable
from celery import Celery
celery_app = Celery("app", broker=settings.celery_broker_url)

Compose / process list:

text
api | worker | beat (1) | flower (1, private) | rabbitmq | redis | db

---

4. Essential configuration

Flag / settingPurpose
--port=5555HTTP port (behind internal proxy)
--basic_auth=user:passBuilt-in basic auth (or put auth at proxy/SSO)
--broker_api=Optional RabbitMQ management API URL for more broker insight
--persistent=TruePersist task state to Flower’s DB (optional; know disk use)
--db=flower.dbPath when persistent
--max_tasks=10000Cap history size
--xheadersTrust proxy headers when TLS terminates upstream
--url_prefix=flowerIf mounted under a path
bash
# PSEUDOCODE : richer RabbitMQ view (management plugin + credentials)
celery -A app.workers.celery_app.celery_app flower \
 --broker_api=https://user:pass@rabbitmq:15672/api/

Environment variables (common):

bash
CELERY_BROKER_URL=amqps://...
FLOWER_BASIC_AUTH=ops_user:strong_password
FLOWER_PORT=5555

---

5. Security (mandatory)

Never put Flower on the public internet without controls.

ControlRequirement
NetworkPrivate VPC / cluster network / VPN / mesh only
AuthBasic auth or SSO at reverse proxy (OAuth2 proxy, etc.)
TLSTerminate TLS at ingress/proxy
AuthorizationOps-only roles; not all developers by default
ActionsDisable or restrict revoke/shutdown if your policy requires
SecretsFlower creds ≠ DB creds; rotate
nginx
# PSEUDOCODE : internal reverse proxy sketch
# listen only on internal LB
location /flower/ {
 auth_request /oauth2/auth; # or htpasswd
 proxy_pass http://flower:5555/;
 proxy_set_header Host $host;
 proxy_set_header X-Forwarded-Proto https;
}

Same rules for RabbitMQ Management UI.

---

6. What to watch in the UI

Workers tab

SignalMeaning
Worker missingProcess crash, deploy, wrong broker URL
ConcurrencyPrefork pool size vs load
QueuesIs worker bound to jobs.io / jobs.heavy?

Tasks

StateMeaning
ActiveRunning now
ReservedPrefetched (with prefetch_multiplier=1 this stays small)
Succeeded / FailedNeeds task events for reliable history

During incidents

  1. Any workers online?
  2. Active task stuck for too long? (compare to soft_time_limit)
  3. Failure rate jumping?
  4. Then check RabbitMQ depth/DLQ and app metrics : not Flower alone

---

7. Enable Celery events

Flower relies on Celery events for live task visibility.

python
# PSEUDOCODE : celery config (workers)
celery_app.conf.update(
 worker_send_task_events=True, # workers emit events
 task_send_sent_event=True, # optional: task-sent events
)
bash
# workers must not disable events
celery -A app.workers.celery_app.celery_app worker -E -Q jobs.default,jobs.io
# -E is --task-events

Without events, Flower may show workers but sparse/empty task history.

---

8. Flower + metrics (better together)

NeedFlowerMetrics / logs
“Who is online?”ExcellentWorker up gauge
Queue depth / DLQWeak / secondaryRabbitMQ exporter (primary)
p95 time-to-completeApproximatejobs table + histograms
AlertingNot an alerterAlertmanager / PagerDuty
Trace one job_idSearch if presentStructured logs + OTel

Minimum production set:

  1. Flower (private) for humans
  2. Prometheus metrics: enqueue, success, fail, runtime, depth, consumers
  3. Alerts: zero consumers, DLQ > 0, success drop, SLO breach
  4. JSON logs with job_id + task_id

---

9. Deploy sketch

yaml
# PSEUDOCODE : docker-compose fragment
services:
 flower:
 image: your-app:${TAG}
 command: >
 celery -A app.workers.celery_app.celery_app flower
 --port=5555
 --basic_auth=${FLOWER_BASIC_AUTH}
 environment:
 CELERY_BROKER_URL: ${CELERY_BROKER_URL}
 # no public ports in production : only attach to internal network
 networks: [internal]
 depends_on: [rabbitmq]
 restart: unless-stopped

Kubernetes:

  • Deployment replicas: 1 is enough for Flower
  • Service ClusterIP only
  • Ingress with auth + TLS, or no Ingress (port-forward / VPN)
  • Resource limits modest (CPU/memory grow with --persistent and task volume)

---

10. Troubleshooting

SymptomChecks
Empty workersBroker URL; workers running; same vhost; network policy
Empty tasks-E / worker_send_task_events; clock skew
Flower OOMLower --max_tasks; disable or bound persistence
Stale stateRestart Flower; check broker connectivity
Auth loopsurl_prefix, proxy headers, cookie paths
“It works in Flower but user status wrong”You’re using Flower as product truth : fix API/DB

---

11. Checklist

  • [ ] Flower runs as its own process/container (not inside API)
  • [ ] Same Celery app module / broker as workers
  • [ ] Task events enabled on workers (-E / config)
  • [ ] Private network only
  • [ ] Authentication enabled (basic or SSO)
  • [ ] TLS at the edge
  • [ ] Documented: Flower ≠ user job status
  • [ ] RabbitMQ depth/DLQ/consumer metrics + alerts still configured
  • [ ] Optional --broker_api only over TLS with locked-down creds
  • [ ] Revoke/admin actions limited by policy
  • [ ] Runbook link from dashboard (scale workers, redrive DLQ)

---

Quick commands

bash
# start
celery -A app.workers.celery_app.celery_app flower --port=5555 --basic_auth=user:pass

# worker with events
celery -A app.workers.celery_app.celery_app worker -E -Q jobs.default,jobs.io -c 4

# never: expose 5555 on 0.0.0.0 to the internet without auth

---

Prometheus: 19 Integrate Flower with Prometheus

See also

External