chore: reference for user roles and for worker operations

This commit is contained in:
Björn Fromme
2026-08-12 17:55:37 +02:00
parent 3b116add75
commit 44fab8bf6c
2 changed files with 472 additions and 0 deletions
+135
View File
@@ -0,0 +1,135 @@
# Operations: messenger workers and cron
Reference for how queued work actually gets executed in production. Describes the setup as
it is — not a plan.
The governing fact, from which most of the rest follows:
> **Nothing in this repository starts a worker.** Both workers are crontab entries on the
> host, outside version control. `deploy.php` only ever *stops* workers, and even that is
> largely ineffective (see below). If you change how a queue is consumed, change it here
> too, or the next person has to reverse-engineer it from the config.
---
## Transports
Defined in `config/packages/messenger.yaml`, all backed by the same Doctrine table
(`MESSENGER_TRANSPORT_DSN=doctrine://default?auto_setup=0`):
| Transport | `queue_name` | Carries | Consumed by |
|-----------|--------------|---------|-------------|
| `async` | `default` | all ordinary mail, chat and SMS notifications | the async cron worker |
| `mailing` | `mailing` | bulk teamer mailings only | the mailing cron worker |
| `failed` | `failed` | anything that exhausted its retries | nobody — inspected by hand |
| `sync` | — | dev and test only | the request itself |
The two live workers are isolated **by the `queue_name` column**, not by how they are
started: the Doctrine transport's `get()` filters on it, so neither worker can see the
other's messages. Ordinary mail therefore never queues behind a mailing of several hundred
recipients.
Overlapping runs of the same worker are harmless. `Connection::get()` selects
`FOR UPDATE SKIP LOCKED` inside a transaction and stamps `delivered_at` before committing,
so two consumers can never be handed the same row. No `flock` is needed; an overlap only
costs a second process.
## The cron entries
> **The async line below is a placeholder.** Read the real one off the host
> (`crontab -l` as `p704161`) and transcribe it verbatim, including any differences in
> flags, php path or logging — those differences are information.
```cron
# ordinary mail, notifications
*/15 * * * * cd /home/www/p704161/html/myep-team/current && /usr/local/bin/php bin/console messenger:consume async --time-limit=... --memory-limit=... --quiet >> /home/www/p704161/html/myep-team/shared/var/log/cron-messenger-async.log 2>&1
# bulk teamer mailings
*/15 * * * * cd /home/www/p704161/html/myep-team/current && /usr/local/bin/php bin/console messenger:consume mailing --time-limit=300 --memory-limit=128M --quiet >> /home/www/p704161/html/myep-team/shared/var/log/cron-messenger-mailing.log 2>&1
```
Why each part is the way it is:
- **`/usr/local/bin/php`, not bare `php`.** `deploy.php` sets `bin/php` to this path
explicitly for both hosts, because the host's default CLI is not the right one. Cron's
`PATH` is typically only `/usr/bin:/bin`. `vendor/` is rsynced from a developer machine
and there is no `deploy:vendors`, so the CLI must match the PHP the dependencies were
installed against (8.3).
- **`--time-limit=300`** on the mailing worker. An 800-recipient mailing needs roughly
200360 seconds of worker time — the fan-out renders 800 Twig bodies and inserts 800
rows, then the mails go out at about 35 per second over a reused SMTP connection. So a
mailing usually drains in one window and occasionally spills into the next; worst case
end to end is about half an hour. Raise it toward 840 for lower latency at the cost of an
effectively always-on process, lower it for a lighter footprint. Anything under 900
guarantees two runs cannot overlap.
- **Stopping mid-mailing is safe.** The time limit is only checked *between* messages
(`StopWorkerOnTimeLimitListener` listens on `WorkerRunningEvent`), so it never truncates
a send or a half-finished fan-out. Unhandled rows keep `delivered_at = NULL` and the next
run picks them up.
- **`--memory-limit` is a recycle trigger, not a safety net.** It too is only checked
between messages, so it cannot protect the fan-out, which is the memory peak. That is
down to the CLI `memory_limit` ini — keep it at 256M or more.
- **`--quiet` plus a redirect.** `messenger:consume` prints a banner on every run and
unredirected cron output is mailed to the crontab user, which would be 96 mails a day.
Nothing is lost: errors still reach `var/log/framework.prod.log` through the prod monolog
handler. Redirect into the **shared** `var/log`, which is in `shared_dirs` and survives
deploys — never anywhere under `var/cache`, which is per-release by design.
- **No `APP_ENV=prod` on the line.** `bin/console` loads Dotenv relative to the project
dir, and `APP_ENV` lives in the shared `.env.local`. A real environment variable takes
precedence over that file, so putting it in the crontab would create a second source of
truth and would need a different line for staging, which runs on the same host under the
same user.
## Deploys do not restart workers
`deploy.php` ends with `deploy:stop-workers`, which runs `messenger:stop-workers`. That
writes a restart timestamp into a cache pool under `%kernel.cache_dir%`, i.e. into the
**new release**. A worker that is already running was started from the previous release and
stays pinned to that release's `var/cache` for its whole life, so it never sees the flag.
**`--time-limit` is therefore what actually bounds a worker's life and what makes new code
take effect** — up to five minutes after a deploy for the mailing worker. This is not worth
"fixing" by moving `var/cache` into `shared_dirs`; that would break cache warming and
opcache invalidation far worse than the problem it solves.
Related: `keep_releases: 3`. A worker must not outlive three deploys, or it will be running
out of a directory that has been deleted. At a five-minute lifetime this cannot happen.
## When something goes wrong
```bash
bin/console messenger:stats # queue depths, including failed
bin/console messenger:failed:show # what failed and why
bin/console messenger:failed:show <id> -vv # the full exception
bin/console messenger:failed:retry # requeue, interactively
```
Nothing watches the `failed` transport, so a failed mailing stays silent until somebody
looks. Check it after any large mailing.
The scheduler's `email_on_failure` does not help here, and covers less than it appears to:
a task only counts as failed when its command exits non-zero, and `App\Command\CronCommand`
always returns `Command::SUCCESS` — even where it renders `$io->error()` for a failed log
flush. So none of the five services it orchestrates can currently raise an alert.
Three failure modes worth knowing because they are quiet rather than loud:
- **`.env.local` missing from the deploy path.** `APP_ENV` falls back to `dev` and
`MAILING_MAILER_DSN` to `null://null`. The worker then consumes every message and
delivers nothing, reporting success and recording no failure. The quietest failure in the
whole system.
- **No worker consuming `mailing`.** Mailings queue up indefinitely while the admin is told
they are on their way. `messenger:stats` shows it immediately.
- **Wrong PHP CLI.** A version mismatch against the rsynced `vendor/` fails at parse time,
inside a cron job nobody reads. This is why the log redirect exists.
## Scheduled tasks are a different mechanism
`config/packages/zenstruck_schedule.yaml` holds the daily jobs (`app:cron`,
`app:bpn-import`, `app:teamer-status`, …), driven by a separate `schedule:run` cron entry.
**Do not put `messenger:consume` in there.** The bundle runs tasks in-process and
sequentially, so a consume task holds the entire schedule for its duration — every later
task waits, and the next `schedule:run` starts concurrently because no task declares
`withoutOverlapping`. This was tried and reverted (`dc796849`, "remove messenger from
scheduler config to be run separately"). Workers belong in their own crontab lines.