.. This file is part of PEPSI. Copyright (C) 2026 GNUnet e.V. PEPSI is free software; you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation; either version 3, or (at your option) any later version. ============ Architecture ============ Every post-ingress processing step is an independent **stage program** driven by a **single database table**. This chapter covers the workspace layout, the data model, the stage contract, the dispatcher and the message lifecycle. Workspace layout ================ Pepsi is a Cargo workspace (edition 2024). The **workspace root is itself the ``pepsi`` package** and owns every ``[[bin]]`` in ``src/bin/*.rs``; the ``pepsi-*`` member crates are **libraries** whose ``run()`` the thin binaries call. So binary ``pepsi-X`` lives in ``src/bin/pepsi-X.rs`` and delegates to crate ``pepsi-x``'s library. * ``pepsi-common`` — shared infrastructure: domain validation (``domain``), DKIM/ARC keys and signing (``keys``/``sign``/``auth``), the outbound SMTP and LMTP client (``smtp``), the SRS engine (``srs``), DSN types (``dsn``), the MIME downgrade machinery (``mime``), the header/body split (``message``), the listening-socket layer (``net``), the local-recipient resolver (``local``), the per-address override layer (``settings``), the RFC 3834 auto-reply rules (``autoreply``), and the **stage scaffolding** (``stage``). * ``pepsi-ingress``, ``pepsi-dispatch``, ``pepsi-httpd``, ``pepsi-setup``, ``pepsi-keydisc``, ``pepsi-sendmail``, ``pepsi-telemetry-client``, ``pepsi-telemetry`` — the non-stage services and clients. * ``pepsi-queue``, ``pepsi-status``, ``pepsi-config``, ``pepsi-settings``, ``pepsi-whitelist``, ``pepsi-quota``, ``pepsi-keys``, ``pepsi-tlsrpt``, ``pepsi-failure-bouncer``, ``pepsi-list``, ``pepsi-archive``, ``pepsi-secure-link`` — the operator command-line tools (none of them a stage; the last three also hold the mailing-list, archive and secure-link logic ``pepsi-httpd`` serves). * ``pepsi-crypto``, ``pepsi-keymat``, ``pepsi-setup-model``, ``pepsi-stage-validate`` — further shared libraries (the OpenPGP/S/MIME protocol code, stored-key loading, the setup interview model, and the per-stage configuration validators). * ``pepsi-helper-*`` — the small privileged or single-purpose helpers a stage or a tool ``exec``\ s (``maildir-writer``, ``dot-forward``, ``auto-pay``, ``mailbox-scan``, ``token-refresh``). * ``pepsi-stage-*`` — the stage programs. * ``vendor/taler-rust`` — a vendored git submodule (a *separate* workspace, excluded here) providing ``taler-common``'s config/CLI/database machinery (``taler_main``, ``ConfigSource``, ``Config``/``Section``, ``taler_common::db``). Each program defines a ``constants::CONFIG_SOURCE`` identifying it (project/component/exec), and stage programs additionally define a ``constants::PROGRAM`` — the binary name used as ``PROGRAM`` in a stage section. The unified binary ================== A development build (the default ``multibin`` Cargo feature) compiles each program to its own executable. A **release** build instead enables the ``unibin`` feature, which links most programs into one multi-call executable, ``pepsi``: a small ``main`` inspects ``argv[0]`` and dispatches to the matching program's entry point. ``make install`` places that single binary in ``$PREFIX/libexec/pepsi/pepsi`` and creates a ``$PREFIX/bin`` symlink per program name pointing at it (``pepsi-ingress -> ../libexec/pepsi/pepsi``, and so on). Operators, the dispatcher and the systemd units keep using the per-program names unchanged; only the on-disk shape differs. The folded programs' ``main``\ s live in the root package's own library (``pepsi::programs::``), so the same code backs both the standalone binaries and the unified one. The motivation mirrors why a C project prefers a shared library over statically linking the same code into every executable. Pepsi's programs share a great deal of code — the async runtime (Tokio), the SQL and TLS stacks (sqlx, rustls), the SMTP/MIME/DKIM machinery in ``pepsi-common``, the CLI/config layer, and more. When each program is its own statically-linked binary, every one carries a private copy of all that common code. Folding them into a single binary keeps **one** copy, with three intended benefits: Smaller on disk The forty-odd folded programs (``FOLDED_BINARIES`` in the ``Makefile``) become one binary plus a symlink (a few bytes) each, instead of forty multi-megabyte executables that each re-embed the shared code. Faster process start Pepsi is process-heavy: the dispatcher runs each stage as a *pool* of worker processes and re-spawns them (after ``MAX_MESSAGES``, a timeout or a crash; see `The dispatcher`_). Because every worker — of *every* stage — ``exec``\ s the **same** file, that file's text is read from disk at most once and then served from the kernel page cache for every subsequent spawn. With separate binaries, the first spawn of each distinct stage pays its own cold-cache read. Less physical memory The read-only segments (code and read-only data) of an ``mmap``\ ed executable are backed by shared physical pages: every process running the file maps the *same* physical copy (this is independent of ASLR, which only randomises the virtual address, not the page-cache backing). With one binary the dispatcher, ingress, the HTTP server and all the stage workers — potentially dozens of processes — share a single resident copy of the common code. With separate binaries, each distinct program that is running keeps its own resident copy of that duplicated code, exactly the redundancy a shared library removes. Two classes of program are deliberately **kept separate**, installed as their own ``$PREFIX/bin`` executables (their ``[[bin]]`` targets build in both feature modes and are excluded from the unified binary): * **Programs deployed or invoked separately.** ``pepsi-stage-detect-language`` links the ``lingua`` language models — a large dependency nothing else uses — and ``pepsi-setup`` pulls in an HTTP client and the one-off provisioning logic. Folding either in would *grow* the shared binary, and every process's resident set, with code only one program uses, so they stay out. ``pepsi-config`` is likewise left standalone; ``pepsi-telemetry`` ships in its own Debian package and runs on another host; and ``pepsi-helper-mailbox-scan`` carries **no** permission bit but is ``exec``\ ed by name from ``pepsi-whitelist``. (``pepsi-stage-detect-language`` is itself a second, two-program multi-call binary: the same file is also :manpage:`pepsi-detect-language(1)`, installed as a ``$PREFIX/bin`` symlink to it.) * **Privileged binaries.** Eleven programs carry their own permission bit: the setuid-root helpers ``pepsi-helper-maildir-writer``, ``pepsi-helper-dot-forward`` and ``pepsi-helper-auto-pay``; their setgid stage callers ``pepsi-stage-relay-to-maildir``, ``pepsi-stage-dot-forward``, ``pepsi-stage-auto-pay`` and ``pepsi-quota``; the setgid ``pepsi-stage-relay-to-smarthost`` (which reads the ``pepsi-token`` group's OAuth token files); the setuid ``pepsi-whitelist``; and ``pepsi-stage-encrypt`` / ``pepsi-stage-decrypt``, setuid ``pepsi-crypto`` because PostgreSQL peer authentication keys off the *effective uid* and that role alone is granted ``crypto_identity.private_wrapped``. A symlink cannot hold a setuid/setgid bit — the bit must live on the target file — so folding these in would force that bit onto the shared binary, where it would apply to *every* symlinked program and collapse the privilege separation described in :doc:`installation`. (``pepsi-keys`` is standalone for the same family of reasons but ships **without** a bit; ``debian/pepsi.postinst`` is the authoritative list of what is actually set.) .. note:: The start-time and memory wins are realised at runtime by the long-running, process-heavy services (the dispatcher and its worker pools, where many processes run the one file concurrently); a one-shot operator command invoked once — say ``pepsi-status`` over SSH — sees the disk-size win but no start-time benefit from sharing. The standalone binaries (``STANDALONE_BINARIES`` in the ``Makefile`` is the authoritative list) also still each embed their own copy of the common code, so the de-duplication is large but not total — a deliberate trade for the two reasons above. The data model ============== One PostgreSQL schema, ``pepsi``, holds 61 tables. Two of them are what the whole system turns on: the **message queue**, and the **configuration overlay**, which lives there because the database is the only shared, writable, transactional state a deployment has. Most of the rest are companions to one of those two — caches, counters and per-address overrides the pipeline consults. The remainder belong to self-contained subsystems that keep their own state in the same schema: the end-to-end key store (below), the secure-link portal, the confirm-to-send secretary, mailbox quotas, the administrative surface, the mailing lists (23 ``list_*`` tables and ``mailing_list``) and their web archive (nine ``archive_*`` tables). ``pepsi-setup/db/pepsi-0001.sql`` is the authoritative list. ``pepsi.workqueue`` One row per message in flight. Besides the envelope and parsed metadata, the message itself is stored **split** into a ``headers`` column and a body (split at the first blank line; the invariant is ``raw = headers || CRLF || body``). Header-only stages never load or rewrite the body. The body is a row of ``pepsi.workqueue_body``, referenced by ``body_id`` and **shared** by every row split off the same message (per-address settings, a relay stage's one row per recipient, a list fan-out, bounce and DSN clones), so a message to *N* recipients is stored once, not *N* times. A stage that rewrites a body gets a fresh ``workqueue_body`` row for its own message only (copy-on-write, in ``commit_row``), so a sibling never sees another's rewrite. Bodies are released by a trigger when the last row referencing them goes, and **pepsi-dispatch** sweeps what a race leaves (``workqueue_body_gc``). The pipeline-control columns are: * ``stage`` — the ``[stage-]`` section of the program currently responsible for the message (``init`` for a fresh message). * ``status`` — ``pending`` / ``running`` / ``paused`` / ``failed`` / ``timeout`` (a PostgreSQL ``ENUM``). * ``state`` — free-form ``JSONB`` carried with the message (see `The state object`_). * ``timeout`` — when a ``paused`` message becomes due for retry. The authentication verdicts (``spf``/``dkim``/``dmarc``/``arc``) and the ``authserv_id`` live in ``state`` (under the ``auth`` and ``origin`` keys respectively), not in dedicated columns. ``pepsi.dns_address`` A cache of resolved MX-host A/AAAA addresses with per-address connect health, maintained by ``pepsi-stage-relay-to-internet``. A working address is preferred until its DNS TTL elapses; a host is re-resolved once all its addresses have failed or expired. ``pepsi.stage_stats`` / ``pepsi.dispatch_stats`` Cumulative pipeline statistics: per-stage message counts, total processing time, kills/timeouts and crashes, plus the global stage/message totals. The global row also carries ``messages_failed`` (messages that ended in a terminal ``failed``/``timeout`` state) and ``serialization_failures`` (transient ``40001``/``40P01`` serialization/deadlock retry events the dispatcher observed) — both expected to stay at zero under normal operation, so a rising count signals a defect rather than a load limit. ``pepsi-dispatch`` is the **only** writer: it accumulates the deltas in memory and flushes them in one transaction roughly once a minute, whenever the pipeline goes idle, and on shutdown. Stages fused into another stage's worker pass are reported back to it on that pass's worker status line rather than written by the worker, so that a table with one row per stage never has every worker of a stage contending for it. ``pepsi-httpd``'s ``/metrics`` reads them (and the live active/paused gauges straight from ``pepsi.workqueue``). ``pepsi.config_override`` The administrator-managed configuration layered on top of the INI file: one row per ``(scope, section, option)`` override, where *scope* is ``global``, ``domain:`` or ``address:
``. This is what lets a running deployment be reconfigured without editing files. It works the same way the queue does: a trigger issues a ``config_changed`` ``NOTIFY``, and the components react — ``pepsi-dispatch`` retires its stage workers so their replacements read the new values. Two properties are structural rather than conventional. Sections that must exist before a database connection does — and the listener sections, which are a security boundary — are **never** read from here, so a compromised database cannot move a listening socket or retarget the database connection. And write access belongs to a dedicated ``pepsi-config`` PostgreSQL role that no mail-processing component holds: a stage worker may read the configuration and may not change it. See :doc:`configuration`. ``pepsi.settings`` The per-address override layer, and a deliberately *separate* table from ``config_override`` because account owners write it themselves by e-mail. The distinction is a privilege boundary, not duplication. Rows are inserted by the ``workqueue_add`` stored function, which deliberately does **not** notify: PostgreSQL holds a database-wide lock from the moment a transaction queues a notification until it commits, so keeping ``pg_notify`` out of the admission transaction is what lets concurrent SMTP sessions group-commit. ``pepsi-ingress`` coalesces instead and wakes the dispatcher **once** per burst, out of band and after the rows commit (``DISPATCH_WAKE_INTERVAL``); writers *outside* the dispatcher's loop — ``workqueue_inject``, ``workqueue_resume``, ``keydisc_release``, ``pepsi-failure-bouncer`` and the ``pepsi-queue`` repair commands — notify explicitly (through ``pepsi_common::db::notify_workqueue``), and writers inside it stay silent because the worker's own status line brings the coordinator round to the same claim. The contract is written out at the top of ``pepsi-setup/db/procedures.sql``. A stage injects a brand-new *side* message with ``workqueue_inject`` (via ``StageContext::enqueue_new``), and spawns **clones** of the current row — a delay DSN, a per-recipient bounce, a success DSN — with ``workqueue_pause`` or ``workqueue_finish_with_clones``, which clone it server-side so the body never round-trips through the application; the clone's unique ``token`` makes it at-most-once. Five further tables form the end-to-end cryptography **key store**, which has no part in the message state machine and is managed by its own tool (:doc:`programs/pepsi-keys`; the design is in :doc:`key-management`): ``pepsi.crypto_identity`` The key pairs of the addresses we serve — public material plus the **private** half — one row per capability (``sign``, ``encrypt`` or ``both``), with its lifecycle columns (``status``, ``is_primary``, ``published``, ``expires_at``, ``private_purged_at``). ``pepsi.peer_key`` Remote correspondents' cached public keys and certificates, each with the ``source`` it came from, whether that lookup was DNSSEC-validated, the last validity verdict and a cache deadline. ``pepsi.peer_has_own_key`` Correspondents proven to hold one of our own keys (they encrypted to it and signed with their own address), so our key file is not attached to mail for them. Written by ``pepsi-stage-autocrypt-learn``. ``pepsi.ca_trust`` The CA certificates an inbound S/MIME chain must reach to count as trusted. ``pepsi.key_request`` The in-flight key-discovery requests ``pepsi-keydisc`` answers, one row per address, carrying the methods asked and the methods that have replied, and — once settled without a key — doubling as the **negative cache** (``negative_until``) that keeps a second message for the same address from parking again. Seven more tables belong to the **administrative surface** (:doc:`admin-api`) and the online setup, and are likewise outside the message state machine: ``pepsi.admin_account`` / ``pepsi.admin_session`` / ``pepsi.api_token`` Remote administrator accounts (Argon2id password hashes), their live browser sessions, and the bearer tokens automation authenticates with. Sessions are rows rather than process memory, so a restart does not log everyone out and two ``pepsi-httpd`` processes agree about who is logged in. Only the digests of a session cookie and of a token's secret half are stored: a database backup contains no usable credential. ``pepsi.event_log`` The **audit log** — one row per configuration change, key operation, login, failed login, account or token change and administrative queue action, with the principal that performed it. It is written by the API *and* by the operator command-line tools, through one helper in ``pepsi-common``, so it is complete regardless of which surface acted; a log that recorded only what happened over HTTP would invite the wrong conclusion from an absence. ``pepsi.mail_log`` The **opt-in** per-message record (``[pepsi] MAIL_LOG``, off by default). Since a delivered message's row is deleted, an ordinary deployment keeps no per-message log at all. Switching this on makes the deployment keep a record of who corresponds with whom, so it is an explicit choice with a bounded retention. See :ref:`mail-log`. Two more belong to the **online setup** (:doc:`installation`), and carry the sharpest privilege boundary in the schema: ``pepsi.setup_task`` / ``pepsi.setup_task_log`` The **intent queue** the browser-driven setup writes and a root program drains. Setting a mail server up is privileged — writing ``/etc/pepsi/pepsi.conf``, handing each ``secrets.d`` fragment to the one account that reads it, running certbot, creating database roles, generating keys — and ``pepsi-httpd`` drops privileges before it accepts a connection. So it does not act: it writes a row describing *what should be true*, from a closed set of eleven kinds with strictly validated parameters, and ``pepsi-setup apply`` decides how. Root is reached through a table, never through a socket that speaks a protocol, and every request and outcome is a durable row. The second table carries the progress lines a running task emits, so a certbot run can be watched without the applier holding an HTTP connection. A row in this table is a request to a process running as root, so the grants are the tightest in the schema: **no account that processes mail may touch it at all**, ``pepsi-httpd`` may only ``SELECT`` (it enqueues through the separate configuration connection), and only ``pepsi-config`` may ``INSERT``. Which of those wrote a row is not taken on trust: ``written_by`` is forced from ``current_user`` by a ``BEFORE INSERT`` trigger, so it is a fact about the connection rather than a claim in the payload, and the applier refuses anything else. The full trust model is in :manpage:`pepsi-setup(1)`. These carry the schema's other access boundary. The three credential tables are granted to the ``pepsi-httpd`` role **alone** — a component that could read ``admin_session`` could mint itself an administrative session — and the two logs are **append-only** for everything that processes mail: every role may ``INSERT`` (that is what makes the audit log complete) and none may ``UPDATE`` or ``DELETE``. ``pepsi-setup`` re-applies both restrictions on every run, because the blanket ``GRANT … ON ALL TABLES`` it hands the pipeline roles would otherwise undo them, and then *verifies* them against the live server. ``crypto_identity.private_wrapped`` holds each private key sealed with AES-256-GCM under a key-encryption key that lives only in a ``secrets.d`` fragment, and ``pepsi-setup`` grants that **one column** to the ``pepsi-crypto`` role alone: every ordinary service role (``pepsi``, ``pepsi-ingress``, ``pepsi-httpd``, ``pepsi-telemetry``) has its table-level grant replaced by a column-level one that omits it. So a compromise of an unrelated stage yields the public half of every identity and nothing more, and a stolen database yields ciphertext even for that column. The schema is a numbered patch series with a re-creatable procedures file. A released patch is never changed: the first release, 0.0.0, shipped ``pepsi-0001.sql``, and each later release that changes the schema adds the next patch, which migrates an existing database forward in place. See :ref:`upgrading` and :doc:`extending`. The stage pipeline ================== Everything after ingress is a **state machine over a single table**. The moving parts are: * **Stage sections.** Each ``[stage-]`` names a ``PROGRAM`` and optional ``NEXT_STAGE`` / ``BOUNCE_STAGE`` (:ref:`config-pipeline`), plus the worker-pool knobs ``PARALLELISM``, ``MAX_MESSAGES`` and ``QUEUE_LIMIT`` (how many messages the dispatcher pipelines to one worker at once). Stage names are operator-chosen labels independent of crate names, so the same binary can serve several stages. * **Stage workers.** Each stage runs as a pool of persistent worker processes (``PROGRAM worker``). A worker reads ``workqueue_id``\ s from standard input (pipelined up to ``QUEUE_LIMIT`` at a time, processed strictly in order), does its work per message, calls **exactly one** terminal helper, and writes one status line (``0`` on success) per message to standard output. A status line may carry a second field, a JSON report of the stages **fused** into that pass, which the dispatcher folds into its statistics. For manual operation or tests a single id can be piped to a worker (``echo | PROGRAM worker``); there is no positional one-shot form. * **The dispatcher.** The single long-lived coordinator that claims rows and feeds them to workers. The stage contract ================== Each worker opens the database pool once at start-up, then for every message id calls ``pepsi_common::stage::prepare_on(pool, PROGRAM, id, cfg, load)``. This: #. loads the ``running`` row in a **single** ``SELECT`` (refusing to act unless the row is ``running``), #. resolves the message's ``[stage-]`` section, and #. verifies that section's ``PROGRAM`` matches this binary. The ``load`` argument (``stage::Load``) is **fixed per stage** and selects one of three prepared statements, so the row is fetched in one round-trip pulling only the columns the stage needs of the two potentially-large ones (``headers``, and the body, joined in from ``workqueue_body``); the cheap envelope/status columns are always loaded. ``Load::Metadata`` loads neither big column (envelope/``state`` only — SRS, discard, ``if``, ``aliases``, ``route``); ``Load::Headers`` adds the header block but not the body (bounce, check-whitelist, vacation, block-language); ``Load::Full`` adds both, reassembled with ``ctx.raw_message()`` for stages that transmit or hash the message (relay, ARC, DKIM-sign, detect-language, the crypto stages). The header block, when loaded, is read via ``ctx.headers()``; both it and ``raw_message()`` error if the stage under-declared its ``Load``. Standard output is reserved for the worker status protocol, so a stage body must never write to it. The shared pool runs ``READ COMMITTED`` — at most one writer ever touches a given queue row (the single dispatcher claims serially, a worker mutates only its own claimed row), so serializable isolation buys nothing — yet the terminal helpers and the dispatcher still route their queue writes through ``pepsi_common::db::with_retry`` as cheap insurance against a transient serialization/deadlock failure. A stage that rewrites the message records its changes through the content setters (``set_headers``, ``set_mail_from``, ``merge_state``, …) rather than issuing SQL; the worker folds those changes **and** the stage transition into **one** ``UPDATE``, built from exactly the columns the stage touched (signing stages, for example, change only the ``headers`` column, leaving the body untouched). Because content changes ride the in-memory row, a fused successor sees them and they are persisted with the chain's single commit — fusion never drops an update. The program then calls exactly **one** terminal helper, whose database semantics are the contract: ``advance()`` / ``advance_to()`` Move ``stage`` **and set the row ``pending``** there, so the next stage's worker pool claims it. The stage owns this transition end to end: the dispatcher's only forward-progress write to ``status`` is the claim ``pending``\ →\ ``running``, never the reverse. ``reroute(stage, state)`` Like advance, but merges ``state``; this is how a message reaches a bounce stage. ``pause(state, secs)`` Set ``paused`` and a retry ``timeout`` (the dispatcher re-queues it). ``pause_with(merge, secs, keep_indices, clones)`` The richer pause: it also reduces the row to a recipient subset and/or spawns side **clones** (a delay DSN, success DSNs for the recipients already delivered), pause and clones committing together in one ``workqueue_pause`` call. ``fail(state)`` Terminal ``failed`` (left for an operator). ``finish()`` ``DELETE`` the row (delivered or dropped). ``finish_with(clones)`` deletes **and** spawns side clones (per-recipient bounces, success DSNs) in one ``workqueue_finish_with_clones`` call. ``advance_or_finish()`` / ``complete_success(...)`` The delivery-stage terminal paths: advance to ``NEXT_STAGE`` else finish; and the success-DSN gate that reroutes to ``BOUNCE_STAGE`` when ``ORIGINATE_SUCCESS_DSN`` + ``NOTIFY=SUCCESS``. Three further helpers **fan the row out** rather than concluding it. They reduce this row's recipient set and create sibling ``pending`` rows in one ``workqueue_split`` call, cloning sender/headers/body server-side, and the caller still has to finish this row with one of the terminals above: ``split_off_recipients`` peels an index subset onto a single sibling (local delivery keeping the local recipients and sending the rest on); ``split_to_one_per_row`` peels recipients ``1..`` each onto their own row at the *current* stage (the SMTP relays, which deliver one envelope recipient per attempt); and ``fan_out`` is the general form, whose ``FanGroup``\ s may carry brand-new addresses, so it expresses recipient *rewrites* (``pepsi-stage-dot-forward``, ``pepsi-stage-relay-to-lmtp``). A group is created ``pending`` unless its ``status`` says ``Retry`` (``paused`` at its stage with a retry time — ``pepsi-stage-milter`` parks a recipient the filter deferred this way) or ``Failed``; those two keep the message's ``received_at``, so a ``MAX_LIFETIME`` measured on them counts from arrival. Because they do not persist the accumulated in-memory content writes, they **reject** a stage that combined a content rewrite with one of them rather than dropping it silently — the same guard ``pause`` carries. ``enqueue_new`` stands outside the transition entirely: it injects a brand-new *side* message (an auto-reply, a payment request) at a named stage via ``workqueue_inject``. ``rewrite_recipients`` is a convenience terminal that writes a wholly new ``rcpt_to`` and ``state`` and advances, all in the worker's single commit; ``pepsi-stage-aliases`` uses it, rebuilding the parallel ``state.dsn.rcpt`` in lockstep. Two optimisations sit under that contract without changing it. **Fusion.** When ``[pepsi] ALLOW_FUSION`` is on (the default) and a stage advances to a successor that is folded into the *same* binary, needs no more columns than are already loaded (its ``Load`` is no greater) and does not refuse through its own ``can_fuse`` gate, the worker runs the successor **in this same process** — skipping both the advance ``UPDATE`` and the successor's load ``SELECT``. The chain's accumulated content writes and the final transition are then committed in one ``UPDATE`` at the end. The registry that makes this possible is process-global and installed only by the unified binary, so fusion is inert in a per-program development build. Each fused hop is reported to the dispatcher on the worker's status line so ``pepsi.stage_stats`` still counts it. **Batching.** A body-free (``Load::Metadata``) worker additionally batches *across* messages: each pass greedily drains the ids already pipelined on its standard input (up to the stage's ``QUEUE_LIMIT``, stopping the instant a read would block), loads the whole batch in **one** ``SELECT … = ANY(…)``, runs each body, then commits the deferred advances/fails/requeues in one arrayed ``UPDATE`` per outcome *shape* — the advance target rides per row, so divergent ``if`` branches still share one statement. The fan-out terminals still commit per message. This takes a saturated body-free stage from two round-trips per message to roughly two per ``QUEUE_LIMIT`` messages. Stages that load the headers or the body keep the one-id-at-a-time loop, since their large columns do not array cheaply. The worker protocol is unchanged either way: one status line per input id, in order. The state object ================ The ``state`` ``JSONB`` is the only channel between stages. The terminal helpers **merge** new keys (``state || $new`` in SQL) so earlier data survives; the one exception is ``pepsi-stage-bounce``, which clears ``state`` because a bounce is a new null-sender message. The stock keys are: ``origin`` SMTP-origin provenance seeded by ingress: client IP, HELO/EHLO, ESMTP/SMTPUTF8 flags, the ``BODY=`` declaration, TLS parameters, listener, reverse-DNS/iprev, and the ``authserv_id`` (which the ARC stage needs to reproduce the AAR identity). Read by the ARC stage and the relay stages' downgrade decision. ``dsn`` The RFC 3461 parameters: message-level ``ret``/``envid`` plus a per-recipient array of ``notify``/``orcpt``. **Every stage must preserve it.** ``bounce`` Written by a delivery/discard stage when routing to ``BOUNCE_STAGE``, consumed by the bounce stage: ``kind`` (``permanent``/``success``/``delay``), ``diagnostic``, ``failed_recipient``, the copied ``notify``/``orcpt``/ ``envid``, and (from the relay stages) the next hop's ``remote_mta``/ ``smtp_code``/``enhanced_status``/``phase``/``reply_text``. ``attempts`` / ``last_error`` / ``delay_sent`` Delivery-stage retry scratch. ``dispatch_error`` Written by the dispatcher when it forces a row to ``failed``/``timeout``. The state object has its own chapter — :doc:`state` — and the exhaustive key reference is :manpage:`pepsi.state(7)`. The dispatcher ============== ``pepsi-dispatch`` is the **single** coordinator (run exactly one per system). It ``LISTEN``\ s on the ``workqueue`` and ``config_changed`` channels (reconnecting with decorrelated backoff) with a ``POLL_INTERVAL`` safety-net heartbeat, and: * **Claims in batches.** On each wake it claims ``pending`` work for every stage with spare capacity in **one** ``UPDATE … RETURNING`` (each stage up to its own remaining capacity), sets the rows ``running`` and hands each id to a worker of that stage over the worker's standard input. This claim is its only forward-progress write to ``status``, and it is the sole writer that makes it, so it needs no ``FOR UPDATE SKIP LOCKED``. A stage that *advances* has already set the row ``pending`` at its successor, so the dispatcher simply claims it again on the next scheduling round for that stage's pool (workers are per-stage, so a message is not chained inside one process, except where the successor is fused into the same pass). * **Shares each stage fairly among senders.** Which pending rows fill a stage's spare capacity is decided like CPU time among processes (CFS/EEVDF): each *sender* — the authenticated account, else the connecting address (IPv6 by /64), held in the generated ``workqueue.sender_key`` — accrues the worker time its messages use at the stage, and the claim prefers the rows of the senders with the least, so a burst from one sender interleaves with everybody else's mail instead of going first. A newcomer starts level with the least-served sender (quiet time banks no credit), leads decay with ``FAIR_HALF_LIFE``, and ``FAIR_FIFO_PERCENT`` of the slots stay oldest-first so no row can starve. The accounting is in the dispatcher's memory only and is simply forgotten on a restart; the selection is still the one claim statement above, fed the accounting as arrays. * **Scales elastically, pipelines work.** Each stage's worker pool grows on demand up to its ``PARALLELISM`` and shrinks when idle: no worker runs until a stage sees work; work is packed onto the fewest workers so a surplus goes cold and is reaped after ``WORKER_IDLE_TIMEOUT``; after ``MAX_MESSAGES`` a worker recycles its child. A worker is fed up to ``QUEUE_LIMIT`` ids at once, lifting a stage's in-flight capacity to ``QUEUE_LIMIT × PARALLELISM`` without adding processes or connections. A stage whose ``PROGRAM`` cannot be started is held off with a short spawn cooldown rather than retried in a tight loop. * **Requeues** ``paused`` rows once their ``timeout`` elapses (it sleeps until the earliest due time, or the heartbeat). * **Recovers.** A worker that reports a non-zero status or whose child crashes → ``failed`` (and is replaced); one that does not answer within ``MAX_RUNTIME`` is killed → ``timeout`` (reason recorded in ``state.dispatch_error``). On start-up it resets orphaned ``running`` rows to ``pending``; on ``SIGINT``/``SIGTERM`` it stops workers and resets their (and any claimed-but-unassigned) rows. The dispatcher never processes message content itself — it only runs the stage workers, which record their own outcome on the row. Message lifecycle (end to end) ============================== #. **Receive.** ``pepsi-ingress`` accepts the SMTP transaction, authenticates the message, prepends the trace and ``Authentication-Results`` headers, and splits and stores it at ``stage = init``, ``status = pending``. The dispatcher is woken out of band by a coalescing task, not from the admission transaction. #. **Dispatch.** ``pepsi-dispatch`` claims the row → ``running`` and hands it to a ``[stage-init]`` worker. #. **Stages.** Each stage advances the message, setting it ``pending`` at the next stage itself; that stage's pool then claims it: e.g. ARC → SRS → DKIM-sign → deliver. #. **Deliver.** The delivery stage transmits the mail. On success it ``finish``\ es (or advances). A transient failure ``pause``\ s with backoff; the dispatcher re-queues it later. #. **Bounce.** A permanent failure ``reroute``\ s to ``BOUNCE_STAGE``; ``pepsi-stage-bounce`` builds an (unsigned) DSN, which is then signed and delivered like any other message. A bounce is never itself bounced. To add a new step anywhere in this flow — say a spam filter or an archival hook — you write a stage program and splice its section into the graph; see :doc:`extending`.