.. This file is part of PEPSI. Copyright (C) 2026 GNUnet e.V. PEPSI is free software; you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation; either version 3, or (at your option) any later version. ================== Supported Features ================== The protocol- and policy-level features Pepsi supports, with the governing RFC for each: the SMTP and authentication machinery, the end-to-end cryptography, and the operational surface around both. The :doc:`rfc-index` cross-references every RFC to the code and configuration; the per-program chapters list which program realises each feature. Inbound SMTP reception ====================== ``pepsi-ingress`` is a standards SMTP server (RFC 5321) with these extensions: * **Transport security.** Cleartext, **STARTTLS** (RFC 3207) upgrade, and **implicit TLS** (RFC 8314, e.g. the submissions port 465) — selected per listener via ``MODE``. Submission listeners (RFC 6409, ``SUBMISSION = yes``) may be bound — see `Message submission`_. * **Message body transfer.** Classic ``DATA`` and **CHUNKING/BDAT** (RFC 3030) for large messages. ``BINARYMIME`` is not advertised, and a ``MAIL FROM`` carrying ``BODY=BINARYMIME`` is refused with ``501 5.5.4``. * **SIZE** advertisement and enforcement (``MAX_MESSAGE_SIZE``; oversize → 552). * **PIPELINING** and **ENHANCEDSTATUSCODES**. * **8BITMIME** (RFC 6152) and **SMTPUTF8** (RFC 6531) — see `Internationalised and 8-bit content`_. * **DSN** (RFC 3461) — see `Delivery Status Notifications`_. * **SMTP AUTH** (RFC 4954) for submission — see `Client authentication`_. * Acceptance policy: mail is accepted only for ``ACCEPTED_DOMAINS`` (others are rejected as relaying with ``550``); the domainless ```` mailbox is always accepted (RFC 5321 §4.5.1); recipients that are valid SRS tokens are reverse-decoded and relayed (see `Sender Rewriting Scheme (SRS)`_). * **Durability:** the client is told ``250`` only after the row is committed; if the database is unavailable the client gets a transient ``451`` and retries. Trace headers ============= Ingress prepends an RFC 5321 §4.4 ``Received:`` header to every message *before* authentication (so the ARC seal covers it). The ``with`` clause records the transmission type — ``SMTP``, ``ESMTP``, ``ESMTPS``, ``ESMTPA`` or ``ESMTPSA`` — per RFC 3848 (the bare ``ESMTPA``, authenticated but not over TLS, is the local UNIX submission socket), and a ``for`` clause names the recipient for single-recipient transactions. Client authentication ===================== Each listener can authenticate the *connecting client* and mark its mail as **locally originated** (which lets it relay onward to any domain and stamps ``ESMTPSA`` in the ``Received:`` trace). A session is authenticated when any of four per-listener mechanisms succeeds: * the peer IP is in the listener's ``MYNETWORKS`` (trusted up front); * an **SMTP AUTH** exchange succeeds (RFC 4954, ``PLAIN``/``LOGIN``), verified against a configured SASL backend (``SASL_TYPE``/``SASL_PATH``; Dovecot's auth-client socket today). ``AUTH`` is advertised and accepted **only** over TLS — a cleartext attempt is refused with ``504 5.5.4``, the code RFC 4954 §4 prescribes for a mechanism that "requires an encryption layer" (§1 deprecates the older ``538``); * a presented **TLS client certificate** matches a configured key/CA pin (``TLS_AUTH_CLIENT``); * the peer of a UNIX socket is identified by ``SO_PEERCRED`` (``AUTH_PEERCRED``, the local submission socket) and its login is mapped to at least one address. This one carries a *username*, so it feeds the same ``USERNAME_MAP`` enforcement a SASL session gets. This is distinct from `Boundary authentication`_ below, which verifies the *message* (SPF/DKIM/DMARC) regardless of how — or whether — the client itself authenticated. See :doc:`programs/pepsi-ingress`. Message submission ================== A listener marked ``SUBMISSION = yes`` behaves as a **Message Submission Agent** (RFC 6409) instead of an MX: * **authentication is mandatory** — an unauthenticated ``MAIL FROM`` is refused with ``530 5.7.0`` (so the flag requires one of the client-auth mechanisms above, and TLS on any listener that is not a local UNIX-domain socket); and * **submission fixups** are applied to each accepted message: a missing ``Date:`` (§8.2) and ``Message-ID:`` (§8.3) are added before the message is stored; and * **submission-identity enforcement** (§6) via ``USERNAME_MAP``: the envelope ``MAIL FROM`` and the ``From:`` header must be addresses the authenticated user is permitted to use (otherwise ``550``), and an allowed ``From:`` that differs from the user's canonical identity gains a ``Sender:`` header naming it (§8.1). The flag has no effect on an MX listener. Boundary authentication ======================= Before storing a message, ingress authenticates it so the eventual receiver can still see the verdict Pepsi reached at the boundary (forwarding will break the original SPF/DKIM signals): * **SPF** (RFC 7208) on the ``MAIL FROM`` identity using the client IP. * **DKIM** signature verification (RFC 6376 / RFC 8463). * **DMARC** evaluation (RFC 7489), with optional SMTP-time enforcement (``DMARC_ENFORCE`` → ``550`` on a definite failure under a ``quarantine``/``reject`` policy). Authentication is otherwise **fail-open**: a DNS or parse error is recorded and the mail accepted. * **iprev / FCrDNS** (RFC 8601) — does the client's PTR forward-confirm to its IP. * An **Authentication-Results** header (RFC 8601) recording all of the above, stamped with the ingress hostname as ``authserv-id``. Authenticated Received Chain (ARC) ================================== The :doc:`programs/pepsi-stage-arc` stage implements **ARC** (RFC 8617). It verifies any inbound ARC chain and seals the message as our ADMD by prepending an ``ARC-Authentication-Results`` / ``ARC-Message-Signature`` / ``ARC-Seal`` set, signed with the single ``ARC_ALGORITHM`` (ARC permits one signature per hop). This lets a downstream receiver trust the boundary verdict after forwarding breaks SPF/DKIM alignment. Sealing is fail-open. Because ARC preserves an *upstream* sender's authentication across the forwarding hop, it applies only to mail Pepsi **receives**: locally-originated submissions (``state.local_origin``) are skipped — they are authenticated as the author domain by the DKIM-signing stage — so the generated pipeline places ARC on the inbound branch, after the ``state.local_origin`` split. Sender Rewriting Scheme (SRS) ============================= The :doc:`programs/pepsi-stage-srs` stage rewrites the envelope sender into a local address of a Pepsi-controlled ``SRS_DOMAIN`` that HMAC-encodes the original sender (the truncated MAC is base32-encoded, RFC 4648), so **SPF passes at the next hop**. The null sender and an address already in the SRS domain are left unchanged; an already-SRS address is re-signed in the compact ``SRS1`` form. Ingress performs the reverse direction: a bounce returned to a valid SRS address is verified and relayed to the original sender, while a forged or expired token is rejected (``550``). End-to-end encryption and signing ================================= :doc:`programs/pepsi-stage-encrypt` encrypts a locally submitted message to each recipient, in **OpenPGP** (PGP/MIME, RFC 3156) or **S/MIME** (CMS, RFC 8551 / 5083), signed with the **From: author's** own key. The signature goes *inside* the ciphertext. ``ENABLE_PEP`` (on by default) is a preset of defaults modelled on the pretty Easy privacy project: every local sender at a served domain gets an OpenPGP key automatically, a message is signed only when it is encrypted (``SIGN = encrypted-only``), and cleartext mail carries the sender's key in an ``Autocrypt:`` header but no signature. Any option written out explicitly still wins; ``ENABLE_PEP = no`` makes signing opportunistic and turns the automatic key creation off. Every encrypted recipient gets their own ciphertext on their own queue row: there is no recipient-list disclosure, no content key shared between recipients, and ``Bcc`` leakage is structurally impossible. Recipients that share an outcome share a row. Recipient keys come from the :doc:`key-management` store, filled by the discovery layer (WKD, VKS, DANE, LDAP, inbound harvesting). The stage itself performs **no network I/O**: a recipient whose key is not cached pauses the message until a discovery service settles the request, and one with a fresh negative cache entry takes the no-key path at once. What happens to a recipient with no key is policy — ``ON_NO_KEY`` is ``cleartext``, ``secure-link`` or ``bounce``, and defaults from ``ENCRYPT`` rather than carrying one of its own. ``ENCRYPT = required`` never degrades to cleartext, and no cleartext outcome is ever silent — each is logged and recorded per recipient in ``state.crypto.out``, together with the content container actually used and whether it was a downgrade. The stage runs **before** DKIM signing, so DKIM covers the bytes actually transmitted; see :doc:`configuration` for why that ordering is not optional. Server-side decryption and verification ======================================= :doc:`programs/pepsi-stage-decrypt` is the inbound half. For a message addressed to a recipient this host serves, it opens whatever ciphertext the message carries, verifies whatever signature it carries, records the verdict and removes any security indicator the sender forged. The user reads ordinary mail in their usual client, with no plugin involved. The exception is mail encrypted to a key the user registered from their own mail client (a *client-custody* key, whose private half Pepsi never holds): that is passed through unopened, for the client to decrypt. It runs **only for recipients this host serves**, and for anything else does not even try. Decryption rewrites the body, which invalidates the sender's DKIM signature; that is harmless for a message about to be filed in a mailbox here and is not for one being relayed onward. A message with both kinds of recipient is split, so the forwarded copy is the one that arrived. The verdict vocabulary is precise. ``valid`` means the signature is sound **and** the key was anchored — an X.509 chain to a configured CA, or an OpenPGP key from a ranked discovery source. ``valid-untrusted`` means sound with a key that could not be tied to anything, which is the state of most of the world's signed mail and is **never** reported as ``valid``. ``invalid``, ``unverifiable`` and ``none`` complete the set, and where a message carries several signatures over the same content the worst of them wins. (A signature made *outside* the ciphertext is a separate class: it answers for the message only when nothing signs the plaintext, and then never better than ``valid-untrusted`` — see :doc:`programs/pepsi-stage-decrypt`.) Only ``valid`` sets ``state.signature_verified``, which :doc:`programs/pepsi-stage-check-whitelist` gates a whitelist row on. Both failure paths **deliver by default**, matching Pepsi's fail-open posture on inbound SPF, DKIM and DMARC: a message we could not open is delivered still encrypted (the user may hold the key in their own client), and a bad signature is a recorded verdict rather than a delivery failure (mailing lists that rewrite bodies produce them on entirely legitimate mail). Quarantine and bounce routes exist for deployments that want them. The result reaches the user two ways: an ``X-Pepsi-Crypto`` header with a documented grammar, and — by default — ``[decrypted][verified]`` prepended to the ``Subject`` in nesting order, which for most users is the only signal they will ever see. Both depend on the same discipline: **every** ``X-Pepsi-*`` field and every one of our own subject tags is removed from inbound mail *before* ours is added. A private header is forgeable by definition, and that removal is the entire basis on which it can be believed. The key material a message carries — S/MIME signer certificates, ``application/pgp-keys`` parts, ``Autocrypt`` headers — is read *in memory* to check that message's own signature, because most signed mail carries the certificate that signed it. When it does not, the message pauses on a discovery request rather than blocking on a key server, so even the *first* message from a new correspondent gets a real verdict. The decrypt stage **stores none of it**. All peer-key learning is a stage of its own, :doc:`programs/pepsi-stage-autocrypt-learn`, placed *after* the spam verdict — which is what makes Autocrypt Level 1 §5.3's "ignore messages the MUA believes to be spam" implementable at all (``LEARN_FROM_SPAM``, default off). An inbound pipeline without that stage verifies signatures perfectly well and never learns a correspondent's key. What the decrypt stage opens is filed as plaintext — unless the recipient registered their own mail-client key, in which case :doc:`programs/pepsi-stage-reencrypt`, placed immediately before local delivery, seals it to that key once every reading stage is done. A per-recipient policy (``ON_NO_CLIENT_KEY``) can refuse mail for a user who has no such key rather than file it readable. Key management, discovery and publication ========================================= Neither crypto stage performs network I/O of its own. Both draw on a **key store** in the shared schema, filled from two directions: by :doc:`programs/pepsi-keydisc` — one service instance per discovery method, so a slow key server delays nobody — and by :doc:`programs/pepsi-stage-autocrypt-learn` from the mail that arrives. A stage that needs a key it has not got commits what it may keep, pauses the message and enqueues a request; the service that answers releases every message parked on that address in the same round-trip. * **Discovery sources**, each a separate concurrently-run lookup with its own rank on a trust ladder: the **Web Key Directory** (advanced and direct forms), **DANE** ``OPENPGPKEY`` (RFC 7929) and ``SMIMEA`` (RFC 8162) — accepted only when the resolver's AD bit says the answer was DNSSEC-validated — the **VKS** protocol of a verifying key server, and **LDAP**. Below all of them on the ladder sit the two rungs that are *not* discovery services and have no instance: material **harvested** from inbound mail (S/MIME signer certificates, ``application/pgp-keys`` parts and ``Autocrypt`` headers), and — one rung lower still, on the very bottom — a third party's key introduced by ``Autocrypt-Gossip:`` in a message that arrived encrypted. Both are written inline by :doc:`programs/pepsi-stage-autocrypt-learn`; naming either in ``SOURCES`` is refused. * **Publication** of our own users' keys: the WKD endpoints served by :doc:`programs/pepsi-httpd`, ``OPENPGPKEY``/``SMIMEA`` records printed by :doc:`programs/pepsi-keys` and :doc:`programs/pepsi-setup`, and — only when asked for, per identity — upload to a verifying key server, whose confirmation mail :doc:`programs/pepsi-stage-vks-confirm` answers. * **Custody.** Private material is stored wrapped under a key-encryption key that lives outside the database, and is reachable only by the one database role the crypto stages run as. * A **negative cache** means only the *first* message to or from an unknown correspondent ever waits. :doc:`key-management` is the chapter on all of this; the option reference is :manpage:`pepsi.conf(5)`. The secure-link fallback ======================== Most correspondents publish no key, so ``ENCRYPT = required`` would otherwise leave only "bounce it" or "send it in the clear". The **secure-link portal** is the third answer: :doc:`programs/pepsi-stage-secure-link` stores the message on the server — encrypted under a freshly generated PIN — and mails the recipient a link, while the PIN reaches them by another channel (by default it is mailed to the *sender*, who relays it). The recipient reads the message in a browser, can download its attachments, and, where the operator enables it, reply. The content key is Argon2id over the PIN, a per-message salt and a pepper that lives only in ``secrets.d``, and only the AEAD ciphertext is stored — so a stolen database yields nothing, a stolen server yields nothing, and **no administrator can read a stored message**. The subject travels inside the ciphertext with the body; the sender, recipient, timestamps and lifecycle counters are in the clear because delivery, expiry and "did this reach them?" need them. A lost PIN is a lost message, and the recovery path is that the sender re-sends. Messages are rendered as **text, never as HTML**, so there is no sanitiser to be bypassed and no remote content to phone home; a wrong PIN is bounded by a per-token lockout and a per-source rate limit; and a reply is injected without ``state.local_origin``, addressed only to the stored sender, so the public form cannot become an open relay. See :ref:`secure-link` for the flow, the threat model and the configuration. Outbound DKIM signing ===================== The :doc:`programs/pepsi-stage-dkim-sign` stage prepends **DKIM** signatures (RFC 6376) — both an **RSA-2048** and an **Ed25519** (RFC 8463) signature by default, or either alone under ``[pepsi] DKIM_ALGORITHMS`` — under the ``SIGNING_DOMAIN`` or the message's ``From:`` domain. Gmail, Microsoft 365 and Yahoo do not verify Ed25519, and Google lists such a signature as ``fail`` in its DMARC aggregate reports; DMARC still passes on the RSA signature. The signature always covers the whole body: RFC 6376's ``l=`` body-length tag is never emitted, since strict verifiers reject such a signature outright rather than tolerating the shorter coverage. Signing is fail-closed: a message whose signing domain has no keys, or that cannot be signed, is marked ``failed`` rather than sent on unsigned. Because header selection is bottom-up, the signature does not disturb existing signatures. Delivery Status Notifications ============================= Pepsi implements **DSN** end to end (RFC 3461 / RFC 3463 / RFC 3464): * Ingress advertises ``DSN`` and validates/stores ``RET``/``ENVID`` (on ``MAIL FROM``) and ``NOTIFY``/``ORCPT`` (on ``RCPT TO``) under ``state.dsn``. * The relay stages **propagate** those parameters to a next hop that also advertises ``DSN``, and omit them otherwise (RFC 3461 §6 — Pepsi does not become the DSN-responsible relay). * :doc:`programs/pepsi-stage-bounce` emits the report as an RFC 3464 ``multipart/report`` with RFC 3463 enhanced status codes: * a **failure** report when ``NOTIFY`` requests ``FAILURE`` (the default when absent) — ``NOTIFY=NEVER`` drops silently; * a **success** report (``Action: delivered``) only when the global ``ORIGINATE_SUCCESS_DSN`` is set *and* ``NOTIFY=SUCCESS`` was requested; * a **delay** report (``Action: delayed``) when a relay stage's ``DELAY_DSN_AFTER`` elapses on a still-queued message that asked for ``NOTIFY=DELAY`` (sent at most once). A null-sender message (a bounce) is never itself bounced (RFC 5321 §6.1). Internationalised and 8-bit content =================================== Ingress advertises **8BITMIME** (RFC 6152) and **SMTPUTF8** (RFC 6531) and records the ``BODY=`` declaration. Because a next hop's capabilities are unknown until after connection, the **outbound SMTP client decides per hop**: * If the hop supports the extension, Pepsi re-advertises ``BODY=8BITMIME`` / ``SMTPUTF8``. * Otherwise it **downgrades**: 8-bit MIME leaf parts are re-encoded to a 7-bit transfer-encoding (RFC 2045 — quoted-printable for ``text/*``, base64 otherwise), on the evidence of the octets rather than of the declared encoding; UTF-8 header fields are rewritten as RFC 2047 encoded-words, in the three places RFC 2047 §5 allows one (an unstructured field, a display-name phrase, the inside of a comment); a non-ASCII ``Content-Type`` / ``Content-Disposition`` parameter — an attachment's filename, in any MIME part — takes the RFC 2231 form (``filename*=UTF-8''caf%C3%A9.txt``), and a non-ASCII ``boundary=``, which has no encoded form, is replaced by a fresh ASCII one together with the delimiter lines. * The downgraded bytes are then **re-tested**: conversion is best-effort over a message nobody validated, and RFC 6152 §3 forbids offering 8-bit content to a hop that did not advertise ``8BITMIME`` "under any circumstances", so a body that is still 8-bit afterwards is a permanent failure (bounce) instead. * A non-ASCII **address** (envelope, or inside a header address) that cannot be represented to a non-``SMTPUTF8`` hop is a permanent failure (bounce). Body re-encoding necessarily breaks a body-covering DKIM/ARC signature; this is unavoidable (capabilities are late-bound) and rare. Outbound relay ============== Two interchangeable relay stages send mail off-site: * :doc:`programs/pepsi-stage-relay-to-internet` — **direct-to-MX** delivery (RFC 5321 §5): its own ``MX`` lookup with preference ordering and Happy-Eyeballs address selection (cached per address in ``pepsi.dns_address``), implicit MX via address records (§5.1), and **Null MX** handling (RFC 7505). TLS is authenticated by **MTA-STS** (RFC 8461) with certificate identity checks (RFC 6125); ``enforce`` policies require STARTTLS to a listed MX. * :doc:`programs/pepsi-stage-relay-to-smarthost` — relay through a configured **smarthost**, chosen by recipient domain (or a catch-all), with per-MTA transport (``plain``/``tls``/``starttls``), certificate verification and the full **SMTP AUTH** (RFC 4954) suite — ``PLAIN``/``LOGIN``, ``CRAM-MD5``, ``SCRAM-SHA-1``/``-SHA-256`` (with optional ``-PLUS`` channel binding), ``OAUTHBEARER``/``XOAUTH2``, SASL ``EXTERNAL`` (TLS client certificate), ``NTLM``, ``GSSAPI``/Kerberos and the deprecated ``DIGEST-MD5`` — or ``auto`` negotiation. See `Client authentication`_ and the :ref:`RFC index `. Both stages apply the **loop guard**, counting the message's ``Received:`` headers against their own ``MAX_HOP_COUNT`` option — it is a relay-stage setting in each stage's ``[stage-]`` section, not an ingress one, so a message is stopped when it is about to be sent on again rather than when it arrives. Both stages authenticate the next hop's TLS with **DANE** (RFC 7672, ``DANE = off|warn|strict``) — looking up the hop's ``TLSA`` records and matching them against the presented chain, taking precedence over MTA-STS — and, when ``[pepsi-tlsrpt] SEND_REPORTS`` is on (it is off by default), record every outbound TLS session for **TLS Reporting** (RFC 8460); :doc:`programs/pepsi-tlsrpt` ships the daily aggregate reports. They share retry semantics: a transient failure pauses the message with **exponential backoff** (``RETRY_INITIAL``/ ``RETRY_MAX_INTERVAL``/``RETRY_FACTOR``) until ``MAX_LIFETIME``, after which it is bounced or failed; a permanent failure routes to ``BOUNCE_STAGE`` (or marks the row ``failed``). Local delivery ============== For recipients in a local domain, three stages file the mail on the host instead of relaying it (each forwards the recipients it cannot handle to its ``NEXT_STAGE``, so they compose): * :doc:`programs/pepsi-stage-relay-to-maildir` — writes the message directly into each local user's ``Maildir/new/``. Locality is decided by ``LOCAL_DOMAINS`` and ``TARGETS`` (passwd ``uid`` ranges); the privileged write is done by the setuid-root ``pepsi-helper-maildir-writer``, reached through the stage's own ``pepsi-maildir`` set-group-id bit. * :doc:`programs/pepsi-stage-relay-to-lmtp` — hands the message to a local **Mail Delivery Agent** (typically **Dovecot**) over **LMTP** (RFC 2033), which runs each recipient's **Sieve** (RFC 5228) filter as it files the mail. It delivers all recipients in one transaction and routes each one the MDA rejects onward by its RFC 3463 status. Needs no privileged helper (the MDA drops privilege). * :doc:`programs/pepsi-stage-dot-forward` — processes each local user's ``~/.forward`` file (the classic sendmail/Postfix mechanism) via the setuid-root ``pepsi-helper-dot-forward``, which drops to the user before reading it: forwarded addresses restart the pipeline, ``|pipe``/``/file`` directives run as the user, and a recipient with no ``~/.forward`` passes through unchanged. :doc:`programs/pepsi-stage-aliases` complements these by expanding envelope recipients through a Postfix ``virtual(5)``-style map (full-address or ``@domain`` catch-all, expanded transitively) before local delivery. :doc:`programs/pepsi-stage-vacation` answers mail that arrives while a recipient is away, and tags the forwarded copy's subject so they can see on their return which mail was answered for them. It is the one user-filtering action Pepsi implements itself rather than leaving to an MDA's Sieve, because the reply is a new message that has to be signed and relayed — and because deciding *not* to send one is pipeline knowledge: RFC 3834 says never answer a bounce, mailing-list mail, anything already marked automatic, or a service address, and Pepsi adds its own "never answer spam" and a per-correspondent rate limit. Leave dates are ordinary per-address configuration, so the same option is a national holiday at global scope and one person's holiday in their own ``pepsi.settings`` row. :doc:`programs/pepsi-stage-route` complements them differently: instead of delivering, it *chooses* a next hop per recipient from that recipient's domain, splitting the message when its recipients disagree. It is what lets Pepsi front an existing mail system — the domains behind the gateway go to that system, the rest go to the internet — and it is where ``pepsi-setup`` proves no domain this host serves can be routed back into its own ingress. See :doc:`exchange`. Mailing lists and archives ========================== Pepsi's mailing-list subsystem is a **reimplementation of GNU Mailman 3**: its data model, its rule/chain and handler/pipeline architecture, its REST API, its e-mail command vocabulary and its notice-template names are the GNU Mailman project's design, copyright the Free Software Foundation, and files taken from GNU Mailman, Postorius or HyperKitty remain under the GNU General Public License (``vendor/PEPSI-VENDORING.md`` lists them). The acceptance test is that an unmodified ``mailmanclient``, Postorius or HyperKitty drives it. See :doc:`mailing-lists`, :doc:`archives` and :doc:`mailman-api`. * **Five pipeline stages** rather than a daemon of its own: routing, posting, per-member delivery, e-mail commands and bounce processing — so a list post is an ordinary queue row that ARC, DKIM, SRS, TLS and the relay stages already handle. :ref:`list-architecture` is the map. * **Nine addresses per list and no alias file.** Mailman writes Postfix or Exim maps for every list's nine addresses and reloads the MTA; Pepsi routes on the recipient inside its own pipeline, so there is nothing to regenerate and nothing to get out of sync. * **Moderation with upstream's vocabulary:** eighteen rules, four terminal chains (``accept``, ``hold``, ``reject``, ``discard``) and a seventeen-handler pipeline, with the names the REST API reports. * **RFC 2369 ``List-*`` headers** and **RFC 5064 ``Archived-At``** on every post. * **RFC 8058 one-click unsubscribe** — ``List-Unsubscribe-Post`` plus a keyed ``https:`` URI, honoured with a single ``POST``. Upstream does not have it, and it is the one place Pepsi exceeds it. * **Digests in both formats:** the RFC 1153 plain digest and the MIME ``multipart/digest``, per subscriber, with volume/issue numbering, a size threshold and a periodic timer. * **Bounce processing** with upstream's scoring machine, VERP attribution on every copy (so the seventeen heuristic detectors are the fallback rather than the mechanism), optional probes, warnings and eventual unsubscription. * **An archive** — a reimplementation of HyperKitty, keeping its Message-ID hash and URL scheme so existing archive links survive a migration — with full-text search, optional trigram substring search, mbox import and export, and a ``purge`` that leaves a tombstone. * **Migration from Mailman 2.1 and 3:** the 2.1 ``config.pck`` is read with a restricted pickle machine that cannot instantiate a class (so an untrusted pickle is data, not code), and a Mailman 3 site is read over its own REST API. * **No JavaScript on any list or archive page**, and no ``403`` on a public surface: a refusal is the same ``404`` a nonexistent list gets, because a ``403`` discloses existence. Unsolicited mail ================ Pepsi filters *after* it has accepted a message, so everything below decides what happens to mail already in the queue. Three mechanisms cooperate through one verdict, ``state.spam``, and one table, ``pepsi.whitelist``: * **The correspondent whitelist.** :doc:`programs/pepsi-stage-auto-whitelist` records the recipients of mail your users send; :doc:`programs/pepsi-stage-check-whitelist` recognises their replies (optionally only when DKIM, a verified signature or a named ARC sealer vouches for them) and sets ``state.spam = false``. The operator and users can manage it by hand with :doc:`programs/pepsi-whitelist`. * **Pay-to-send** (:doc:`programs/pepsi-stage-anti-spam`): mail from anybody else is held until its sender pays a small GNU Taler amount. * **Confirm-to-send** (:doc:`programs/pepsi-stage-secretary`): mail from anybody else is held until its sender replies once to a challenge. The reply whitelists them, so a correspondent is asked once, ever. Post-queue milters (SpamAssassin, rspamd, ClamAV, …) and the language filter can set or inform the verdict as well; see :doc:`programs/pepsi-stage-milter`. Shared and per-user whitelists ------------------------------ Every one of these stages names its list with ``WHITELIST_NAME``. A name is either **global** (``correspondents``), or belongs to one account (``alice/sent``, the ``/...`` namespace; ``{login}``/``{localpart}`` templates expand to such a name per recipient where a stage supports them). A business usually wants **one list shared by every employee**: a customer who was told to write to a colleague, rather than to the employee who first mailed them, should not be caught by the spam gate. The operator sets it once:: [stage-auto-whitelist] PROGRAM = pepsi-stage-auto-whitelist NEXT_STAGE = dkim-sign WHITELIST_NAME = correspondents [stage-check-whitelist] PROGRAM = pepsi-stage-check-whitelist NEXT_STAGE = anti-spam WHITELIST_NAME = correspondents The operator may name any list, shared or per-user, in every layer it controls: the INI file, ``pepsi.config_override`` at any scope (a different shared list for one department's ``domain:``, say) and ``pepsi.settings`` rows written with :doc:`programs/pepsi-settings`. An account owner who may edit :doc:`programs/pepsi-stage-auto-whitelist` or :doc:`programs/pepsi-stage-secretary` by mail (:doc:`programs/pepsi-stage-edit-settings`) may only move those stages' ``WHITELIST_NAME`` into their own namespace, or back to the operator's list: both stages *write* the list, so a shared or foreign name would let a user fill it with addresses of their choosing. :doc:`programs/pepsi-stage-check-whitelist` only reads, and has no such restriction. .. caution:: The payment gate exempts only the **null envelope sender**, so it also holds mail nobody will pay for. The secure-link portal's link notification and PIN mail carry the original sender as envelope sender; when they come back in to a gated mailbox on the same deployment they are held until the deadline and then rejected, and the portal stops working without an error. Mailing-list postings carry their authors' ``From:`` addresses, which no correspondent list covers, and their ``List-*`` fields suppress the payment request, so each posting is held and rejected with nobody asked to pay. (Vacation notices and the payment requests themselves use the null sender and are not held.) Whitelist those senders in a group ``[stage-check-whitelist] WHITELIST_NAME`` names: your own domains with DKIM required (the default, which a forged ``From:`` cannot satisfy), and each list by its identifier and ARC sealer:: pepsi-whitelist add correspondents '^[^@]+@example\.com$' pepsi-whitelist add-list --sealer lists.example.org \ correspondents users.lists.example.org :doc:`programs/pepsi-stage-anti-spam` lists the details. Pay-to-send and confirm-to-send compared ---------------------------------------- Both hold unknown senders' mail and both send the sender one null-sender auto-reply, in their language. They differ in what the sender has to prove: .. list-table:: :header-rows: 1 :widths: 24 38 38 * - - Pay-to-send - Confirm-to-send * - The sender proves - they will spend money on this message - they can read mail at the address they wrote from * - Stops - bulk mail at any volume, including from working mailboxes - bulk mail from addresses that cannot receive (most of it) * - Does not stop - a sender willing to pay - a spammer with a working mailbox, who can automate the reply * - Deterrent - economic: every message costs - legal only: ``PENALTY`` names what the sender agrees to pay if the mail was unsolicited, which nothing enforces * - Needs - a GNU Taler merchant backend, and a wallet on the sender's side - nothing but a reply Stated plainly: confirm-to-send is the weaker filter. It is still worth running, because it costs a legitimate correspondent one reply and costs a mail server nothing to operate — and the two combine. With both enabled, the secretary runs first and its ``UNCHALLENGEABLE_STAGE`` points at the paywall, so mail that cannot be asked to reply (a list posting, an unauthenticated or forged sender, a sender over the daily challenge cap) is asked to pay instead, while a confirmed sender reaches the paywall with ``state.spam = false`` and passes it untouched. Confirm-to-send is careful about backscatter, since its challenge goes to the envelope sender and spam's envelope sender is usually forged: it challenges only a sender that passed SPF or DMARC (``REQUIRE_AUTHENTICATED``) and whose ``From:`` is that same address, at most ``MAX_CHALLENGES_PER_SENDER`` times a day across all users; it never answers what RFC 3834 says not to; and its challenge quotes nothing of the held message but a short hint of its subject (``SUBJECT_HINT_LENGTH``, eight characters by default). .. _secretary-spam-filters: Spam filters and the secretary ------------------------------ By default spam-scoring milters (SpamAssassin through spamass-milter, rspamd, MIMEDefang, …) stay **ahead** of the whitelist check, and :doc:`programs/pepsi-setup`'s wizard generates that order: .. code-block:: text … → milter-spamassassin → check-whitelist → secretary → local UNCHALLENGEABLE_STAGE = local A message the filter rejects never reaches the secretary, so obvious spam draws no challenge to its forged sender — that is backscatter avoided. What *does* reach the secretary has passed the filter, which is why the wizard sets ``UNCHALLENGEABLE_STAGE`` to the secretary's own ``NEXT_STAGE``: running the filter again would scan twice. The alternative is to move the filter **behind** the secretary, so it scans only mail that could not be challenged: .. code-block:: text … → check-whitelist → secretary → local └─ UNCHALLENGEABLE_STAGE → milter-spamassassin → local The configuration edit, for a filter at ``[stage-milter-spamassassin]``: 1. Take the filter out of the main chain: set the ``NEXT_STAGE`` of the stage that pointed at it to the filter's own ``NEXT_STAGE`` (normally ``check-whitelist``). 2. In ``[stage-secretary]``, set ``UNCHALLENGEABLE_STAGE = milter-spamassassin``. 3. In ``[stage-milter-spamassassin]``, set ``NEXT_STAGE`` to the secretary's ``NEXT_STAGE`` (here ``local``). Then run ``pepsi-setup check`` (or ``run``) to validate the graph. The trade-off cuts both ways: * **For:** the filter's CPU is spent only on the mail that needs it. Whitelisted mail and mail from confirmed senders — the bulk of a personal mailbox — is never scanned. * **Against:** challenges now go out *unscanned*. Spam the filter would have rejected draws a challenge first, which is more backscatter to forged senders (bounded by ``REQUIRE_AUTHENTICATED`` and the daily cap, but not zero); and spam whose sender does confirm — the one kind confirm-to-send cannot stop — is delivered without the filter ever seeing it. Whichever layout you choose, **virus and policy milters stay on the main path** (ahead of the whitelist check): a confirmed or whitelisted correspondent can still send malware, from a compromised account or unknowingly, and a policy filter's verdict applies to everybody. Operational features ==================== * **Single-table pipeline** with crash recovery: orphaned ``running`` rows are reset on dispatcher start-up; in-flight children are reset on shutdown. * **Elastic, pipelined worker pools and watchdog:** each stage runs persistent worker processes started on demand up to its ``PARALLELISM`` and reaped after ``WORKER_IDLE_TIMEOUT``, recycled after ``MAX_MESSAGES``; each worker is fed up to ``QUEUE_LIMIT`` messages at once (in-flight capacity ``QUEUE_LIMIT × PARALLELISM``, at no extra connection cost), and a worker exceeding ``MAX_RUNTIME`` on its head-of-line message is killed (``timeout``) and replaced. * **Queue tooling:** :doc:`programs/pepsi-queue` lists, deletes, re-stages and bulk-unsticks messages. * **Health monitoring:** :doc:`programs/pepsi-status` prints a read-only summary of the queue, stuck messages, cumulative delivery/failure counters and recent outbound TLS outcomes (also as ``--json``), suitable for running over SSH. * **Provisioning and DNS verification:** :doc:`programs/pepsi-setup` installs the schema, generates keys and prints/validates DNS (DKIM, SPF, MTA-STS). * **HTTP server:** :doc:`programs/pepsi-httpd` serves, over one or more TLS (SNI-selected) or plaintext listeners, the MTA-STS policy file (``/.well-known/mta-sts.txt``), a Prometheus ``/metrics`` page (on an ``ADMIN = yes`` listener only — it describes the pipeline), the **Web Key Directory** endpoints that publish this deployment's own identities (``/.well-known/openpgpkey/…``, direct and advanced forms), the :doc:`secure-link` portal, and the **mail autoconfiguration** document below. * **Mail client autoconfiguration** (*draft-ietf-mailmaint-autoconfig*): a client given nothing but ``fred@example.org`` fetches ``https://autoconfig.example.org/mail/config-v1.1.xml`` and configures itself — submission host, port, transport security and authentication, and the same for the mailbox server. The submission half is derived from the running submission listener, so it cannot drift from the server it describes; the mailbox half names whatever MDA the deployment pairs with. Served publicly and without authentication, as the draft requires, because a client must read it *before* it can know how to authenticate. * **Administrative API and console:** on a listener explicitly flagged ``ADMIN = yes`` — and never on one that would carry it in cleartext off the host — the same server exposes the :doc:`admin-api` under ``/api/v1`` (queue, health, configuration, key store, audit and mail logs) and the browser :doc:`web-ui` under ``/ui``, which is a client of exactly that API. Everywhere else those paths answer the same ``404`` an unknown path gets. Authentication is by UNIX socket peer credentials, a session cookie or a bearer token, and every principal carries scopes; all of it is audited. * **Metrics:** ``pepsi-dispatch`` records per-stage throughput, kills/timeouts, crashes and processing time and the global stage/message totals (flushed to the database in one transaction roughly once a minute), exported by ``pepsi-httpd`` alongside live active/paused gauges read from the queue. * **Structured logging** via ``tracing`` at configurable levels. For standards Pepsi does not implement (``BINARYMIME`` among them), see :ref:`rfc-not-yet`. Feature stability ================= Every feature described in this chapter is tracked in a single :doc:`feature-stability` table, which records — per feature — its governing RFC, whether it has an automated test (and of what kind: ``U`` unit, ``I`` integration, ``I/U`` both), whether it has been verified by hand, and how widely it is deployed and exercised according to anonymous, opt-in usage telemetry (:doc:`programs/pepsi-telemetry`). That table is **generated**: ``contrib/update-feature-stability.sh`` merges the hand-maintained registry ``contrib/feature-registry.tsv`` (name, RFC, test and manual-test columns) with a ``pepsi-telemetry`` ``GET /telemetry/report`` (the deployment and usage counts). See :doc:`extending` for how to keep it current when you add a feature, write a test, or refresh the telemetry counts.