.. This file is part of PEPSI. Copyright (C) 2026 GNUnet e.V. PEPSI is free software; you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation; either version 3, or (at your option) any later version. PEPSI is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU Affero General Public License for more details. ================= Security model ================= This chapter says what Pepsi defends against, what it does not, and **what an attacker gets from each thing they might have taken** — so it is as useful after a compromise as before one. Pepsi is an MTA that holds private keys and can read mail, which makes it worth attacking. The design keeps asking one question: *when this component is compromised, what is still out of reach?* Where the answer is "nothing", that is said here rather than left to be discovered. The *claims* — which assets, against which adversaries, in which custody mode — are stated in :doc:`threat-model`; this chapter is the mechanisms behind them. .. contents:: On this page :local: :depth: 2 What Pepsi assumes ================== If one of these is false, the guarantees below do not hold: * **The host is not already compromised.** Nothing here defends against root on the machine. Root can read every key, every secret and every message. * **PostgreSQL authenticates by peer credentials over a local socket.** Each component connects as its own operating-system user, and the database grants that role only what the component needs. A deployment that switches to password authentication with one shared account throws away most of :ref:`sec-db`. * **The operating system enforces file ownership and the setuid bit.** The private-key boundary is a database ``GRANT`` keyed on the effective uid; a ``nosuid`` mount or a shared account collapses it. * **A validating DNS resolver.** Pepsi trusts the AD bit rather than validating DNSSEC itself (see :doc:`programs/pepsi-stage-relay-to-internet`). A lying resolver defeats DANE and downgrades MTA-STS to trust-on-first-use. * **TLS certificate material is not attacker-supplied.** Certbot integration writes into ``/etc/letsencrypt``; an attacker who can write there can impersonate the server. .. _sec-privilege: The privilege split =================== .. rubric:: Upholds :ref:`tm-adv-component` (what each compromised account reaches), :ref:`tm-adv-local` (no role reachable from an ordinary account) and :ref:`tm-adv-db` (what a service role can write). Pepsi is not one program. Each component runs as its own account, and the separation is enforced by the database as much as by the filesystem. The database side draws two kinds of line, and it is worth being exact about both. **The pipeline role is broad on purpose.** ``pepsi`` (the dispatcher and every ordinary stage) and ``pepsi-crypto`` hold one blanket grant: ``SELECT``, ``INSERT``, ``UPDATE`` and ``DELETE`` on every table in the ``pepsi`` schema, use of every sequence, ``EXECUTE`` on every function, and the same as default privileges for objects created later. Every stage reads and writes the queue, and between them the stages use every pipeline table, so least privilege per stage would need an account per stage; the consequence is stated in :doc:`threat-model` rather than hidden. The grant is then taken back where the pipeline has no business: * ``crypto_identity``: ``SELECT``/``INSERT``/``UPDATE`` only on the columns other than ``private_wrapped`` (``DELETE`` has no column granularity and stays); ``pepsi-crypto`` alone holds the column; * ``config_override``: ``SELECT`` only — writing it is the ``pepsi-config`` role's (``pepsi-httpd`` writes through a ``[pepsi-admin] CONFIG_DB`` connection that authenticates as that role); * ``setup_task`` and ``setup_task_log``: nothing; only ``pepsi-config`` may ``INSERT`` a task; * ``secure_message`` and ``secure_access``: ``INSERT``, ``DELETE`` and ``SELECT`` of every column but ``ciphertext`` on ``secure_message``; ``SELECT`` and ``DELETE`` on ``secure_access``; * ``admin_account``, ``admin_session`` and ``api_token``: nothing; * ``event_log`` and ``mail_log``: append-only (no ``UPDATE``/``DELETE``). **The services that are not the pipeline get what they use.** An audit of every statement each can issue — its own crate and every library it links, following calls into stored functions — sets their grants, and ``pepsi-setup`` refuses to finish if PostgreSQL disagrees: .. list-table:: :header-rows: 1 :widths: 18 47 35 * - Role - Granted - Notably not * - ``pepsi-ingress`` - ``INSERT`` into ``workqueue`` and ``workqueue_body`` (through ``workqueue_add``) and ``SELECT`` of the new ids; ``SELECT`` of the two generated size columns (``header_octets``, ``octets``) that the queue admission check sums; ``SELECT`` on ``config_override`` and ``mailbox_quota``; ``SELECT`` of ``mailing_list``'s ``list_name`` and ``mail_host`` (a list's address, for the recipient check); ``INSERT`` on the two append-only logs. - Any other message in the queue, settings, whitelists, keys, anything about a list beyond its address. A parser bug in the process listening on port 25 reaches one table, to add to it. * - ``pepsi-telemetry`` - Its ``telemetry`` table. - Everything else, so a collector sharing a database with a pipeline cannot reach the mail. * - ``pepsi-httpd`` - The administrative credential tables, the portal, the key tables (not ``private_wrapped``), the list and archive tables, ``event_log``/``mail_log`` (``pepsi-httpd prune``, run by ``pepsi-log-prune.timer``, prunes them); on ``workqueue`` the envelope, ``state`` and routing columns, ``INSERT`` (mail the web forms send) and ``DELETE``; on ``workqueue_body`` ``INSERT`` and ``DELETE`` (a body goes with the last message carrying it; the foreign key refuses any other) and the id and size columns only; read-only statistics, ``dns_address``, ``tls_session``, ``mailbox_quota`` and ``telemetry_client`` (the telemetry daemon's liveness row); ``SELECT`` on ``setup_task``. - **A queued message's headers or body** — the console lists envelopes and has no read path to content, nor can it repoint a message's ``body_id`` at another message's body; ``settings``, ``whitelist``, ``vacation_reply``, ``payment_request_reply``, the secretary tables, ``telemetry``, ``origin_nonce``, ``auto_pay_spend``, ``list_post_rate``; writing statistics, quota or the telemetry daemon's liveness row; ``config_override`` and ``setup_task`` except through its ``pepsi-config`` connection. A table a later schema patch adds is not in ``pepsi-ingress``'s or ``pepsi-telemetry``'s reach until a patch grants it: their default privileges are revoked too. The roles *outside* this set are narrower still: ``pepsi-keydisc``, ``pepsi-whitelist`` and ``pepsi-config`` each hold grants on a few named tables only. ``pepsi-crypto`` holds the blanket grant **plus** ``private_wrapped``. The schema owner, ``pepsi-owner``, which ``pepsi-setup`` itself connects as, holds everything by ownership. .. list-table:: :header-rows: 1 :widths: 24 20 56 * - Component - Runs as - What it can reach * - ``pepsi-ingress`` - ``pepsi-ingress`` - ``INSERT`` into the queue and the two reads above; the SRS secret; TLS material. **No** other message, no private key material, no setting, no whitelist. * - ``pepsi-dispatch`` and most stages - ``pepsi`` - The pipeline grant above: the queue, settings, statistics, whitelist, and the secure-link tables without their ``ciphertext``. **Not** ``crypto_identity.private_wrapped``. * - ``pepsi-stage-encrypt`` / ``-decrypt`` - ``pepsi-crypto`` (setuid, mode ``4750 pepsi-crypto:pepsi``) - The ordinary pipeline access every stage has, **plus** the wrapped private key material and the key-encryption key. Nothing else holds both: the schema owner ``pepsi-owner`` can read the column too, but not the key-encryption key. * - ``pepsi-keys`` - the invoking operator - Installed ``0755``, with **no** setuid bit: one binary able to open every private key in a deployment is a far larger prize than ``pepsi-whitelist``'s single table, so by default only ``root`` or the ``pepsi-crypto`` account can use it, and the PostgreSQL grants are the authority. A site that wants users to manage their own identities may deliberately make it setuid ``pepsi-crypto`` (``chmod 4755``), at which point the program's own per-identity ownership rule starts applying — see :doc:`key-management`. * - ``pepsi-httpd`` - ``pepsi-httpd`` - The tables in the second table above — read *and write*, not read-mostly: its handlers change queue routing and key metadata directly. **No** message headers or bodies, no private key material, no settings or whitelists; no ability to write the configuration overlay or enqueue a setup task except through its ``pepsi-config`` connection. * - ``pepsi-keydisc`` - ``pepsi-keydisc`` - The discovery request queue and ``peer_key``. Never a general ``UPDATE`` on ``pepsi.workqueue``. * - ``pepsi-whitelist`` - invoking user (setuid) - ``SELECT``/``INSERT``/``DELETE`` on ``pepsi.whitelist`` only. * - ``pepsi-quota`` - ``pepsi`` (setgid ``pepsi-maildir``, mode ``2550``) - ``pepsi.mailbox_quota``, and — through the group — the setuid-root helper that measures a mailbox. **Mailbox quota keeps its measuring and its refusing apart.** Only ``pepsi-helper-maildir-writer`` can look inside a ``Maildir`` (mode ``0700``; it is the one program that becomes the user), and only members of ``pepsi-maildir`` may run it — a group ``pepsi-ingress`` is deliberately not in and ``pepsi-setup`` warns if anyone adds it to. So the process listening on port 25 can *read a measurement somebody else took* and refuse a recipient on it, and can do nothing else: ``pepsi-setup`` revokes its write access to ``pepsi.mailbox_quota``, leaving ``SELECT``. Both roads to the accounting are closed to it, and neither closure relies on the other. The crypto stages are **setuid, not setgid**: PostgreSQL peer authentication keys off the *effective uid*, so a stage that merely gained a group would still connect as ``pepsi`` and be refused the key column. It has to *be* ``pepsi-crypto``. Every program carrying a bit ---------------------------- .. rubric:: Upholds :ref:`tm-adv-local`: a local account gains no Pepsi role, and each helper acts with the target user's authority only (boundary B4). A multi-call binary cannot carry a per-program setuid or setgid bit, so each of these is installed as its own executable rather than as a symlink into the shared ``pepsi`` binary (the authoritative list is ``STANDALONE_BINARIES`` in the ``Makefile``; the modes are applied by ``make install`` and, on Debian, by ``dpkg-statoverride`` from the ``postinst``). Everything not listed here is unprivileged. .. list-table:: :header-rows: 1 :widths: 38 24 38 * - Program - Owner and mode - Why * - ``pepsi-helper-maildir-writer`` - ``root:pepsi-maildir`` ``4750`` - Becomes the recipient to write into a ``0700`` ``Maildir``. * - ``pepsi-stage-relay-to-maildir`` - ``pepsi:pepsi-maildir`` ``2550`` - The setgid bit is the sole gate on exec'ing that helper, so the mode is ``2550`` and not ``2755``: world-execute would hand every local account ``egid=pepsi-maildir``. The ``pepsi`` user is deliberately **not** a group member and gets the exec right by owning the file. * - ``pepsi-quota`` - ``pepsi:pepsi-maildir`` ``2550`` - Same helper, so a ``reconcile`` sweep can re-measure from a timer without being root. * - ``pepsi-helper-dot-forward`` - ``root:pepsi-forward`` ``4750`` - Drops fully to the user before reading their ``~/.forward``. * - ``pepsi-stage-dot-forward`` - ``pepsi:pepsi-forward`` ``2550`` - The same gate, for that helper. * - ``pepsi-helper-auto-pay`` - ``root:pepsi-wallets`` ``4750`` - Drives the wallet as the account that owns it. * - ``pepsi-stage-auto-pay`` - ``pepsi:pepsi-wallets`` ``2550`` - The same gate, for that helper. * - ``pepsi-stage-relay-to-smarthost`` - ``pepsi:pepsi-token`` ``2550`` - Reads the group-restricted OAuth token files ``pepsi-helper-token-refresh`` writes, the mutual-TLS client key and the Kerberos credential cache. ``2550`` for the same reason as the stages above: the stage accepts ``-c``, so world-execute would let any local account point it at a server of its own and send it those credentials. * - ``pepsi-whitelist`` - ``pepsi-whitelist:pepsi-whitelist`` ``4755`` - World-executable on purpose: any local user may manage their own namespace, and what stops them touching another's is the rule inside the program, not the file mode. * - ``pepsi-stage-encrypt``, ``pepsi-stage-decrypt`` - ``pepsi-crypto:pepsi`` ``4750`` - Peer authentication keys off the effective uid, so the stage has to *be* the role granted ``crypto_identity.private_wrapped``. Group ``pepsi`` and no world-execute: unlike ``pepsi-whitelist`` these are not user commands. * - ``pepsi-keys`` - ``0755``, no bit - See the table above. A site may make it setuid ``pepsi-crypto``. * - ``pepsi-helper-mailbox-scan`` - ``0755``, no bit - The **unprivileged** half of the ``pepsi-whitelist`` pair: it parses hostile mailbox files, so it is spawned with ``setresuid`` to the invoking user — all three ids, so it can never regain the owner. The runtime directory is an ACL, not a shared group --------------------------------------------------- .. rubric:: Upholds :ref:`tm-adv-local` and :ref:`tm-adv-component`: a socket is reachable by the accounts meant to reach it, and creating one grants no unrelated authority. ``/run/pepsi`` holds sockets belonging to several different accounts — ``pepsi-ingress``'s local submission socket, ``pepsi-httpd``'s admin socket, the reverse-proxy socket the front web server connects to, and the applier doorbell — and it is owned by ``root`` with a POSIX ACL granting ``pepsi-ingress`` and ``pepsi-httpd`` write access by name (``debian/pepsi.tmpfiles``). A shared group would not do: it would have to be a group both are in, and every candidate carries something neither should gain — group ``pepsi`` reads the DKIM private keys, group ``pepsi-admin`` *is* the administrative API's access control. ``pepsi-httpd.service`` alone runs with ``pepsi-admin`` as a supplementary group, so that it can hand its admin socket to that group: the group owns no file and grants nothing but administrator status on that socket, which the process serving it has anyway. ``pepsi-ingress`` never gets it. Mode ``0775`` on that directory is therefore load-bearing rather than a widening: on a directory carrying an ACL the group bits **are** the ACL mask, so ``0755`` silently caps every named grant at ``r-x`` and the next restart cannot create its socket. .. _sec-no-runtimedirectory: No unit may claim it as a ``RuntimeDirectory`` ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The ``tmpfiles.d`` snippet is the directory's **only** owner, and that is a rule rather than a preference. ``systemd`` gives a ``RuntimeDirectory=`` to the declaring unit's own ``User=`` and chowns it — *and everything inside it* — on every start. In a directory with four tenants under three accounts, that recursive chown does not merely decide who owns the directory: it rewrites the ownership of sockets other units have already bound. For example, ``pepsi-httpd.socket`` binds ``/run/pepsi/httpd.sock`` as ``0660`` ``root:www-data`` so the front web server can reach it in reverse-proxy mode. A ``RuntimeDirectory=pepsi`` in ``pepsi-httpd.service`` would chown that socket to ``pepsi-httpd:pepsi-httpd`` on every start, the front server would get ``EACCES`` on connect, and **every MTA-STS policy fetch would get a 502** — with every unit ``active`` and nothing logged, because nothing had failed. The same chown would replace the directory's ``root`` ownership and its ACL outright and hand the other tenants' sockets to whichever unit started last, and without ``RuntimeDirectoryPreserve=yes`` a service stopping would delete the directory, live sockets and all. What those units actually needed from ``RuntimeDirectory=`` was one thing: ``/run`` is read-only under ``ProtectSystem=strict``. So ``pepsi-httpd.service`` and ``pepsi-ingress.service`` say ``ReadWritePaths=/run/pepsi`` instead and take their *permission* from the ACL; ``pepsi-setup-apply.service`` needs neither, running as ``root`` without ``ProtectSystem=strict``. The ``.socket`` units keep ``DirectoryMode=0775`` so the directory still appears if one of them binds before ``systemd-tmpfiles`` has run, and nothing deletes it, because nothing claims it. ``tests/check-systemd-units.sh`` (run by ``make check``) fails the build if a ``RuntimeDirectory=pepsi`` reappears in a shipped unit or in the reverse-proxy drop-in ``pepsi-setup`` writes. The front web server's access is granted the same way, and separately. Reaching ``httpd.sock`` needs two permissions: the socket's own ``SocketGroup=``, and traversal of the directory it sits in. The directory's ``other`` ``r-x`` bits would provide the second, but they exist for an unrelated reason (any local account may submit mail through the submission socket), and relying on them would make policy delivery depend on a permission granted for something else. ``pepsi-setup`` therefore writes an explicit grant to ``/etc/tmpfiles.d/pepsi-reverse-proxy.conf`` when it configures reverse-proxy mode, naming the account it actually detected (``www-data`` on Debian, ``nginx``/``apache`` elsewhere) with ``r-x`` and nothing more, and removes it again with the socket drop-in. It is a separate file from the shipped snippet because it is discovered rather than shipped, and conditional rather than always. The telemetry collector is deliberately not a tenant here at all and lives in ``/run/pepsi-collector`` instead — its own directory, shared with no other unit, which is also what keeps it clear of ``/run/pepsi-telemetry``, the telemetry client's. .. _sec-keys: The private-key custody boundary ================================ .. rubric:: Upholds :ref:`tm-gateway` (keys never disclosed to a passive reader or an ordinary stage; boundary B5). Private key material lives in ``crypto_identity.private_wrapped``, and is protected twice over: **By grant.** ``pepsi-setup`` grants that one column to the ``pepsi-crypto`` role and revokes it from every service account, ``pepsi-httpd`` included; apart from ``pepsi-crypto``, only the schema owner ``pepsi-owner`` can read it, by ownership. This is column-level and it is absolute: PostgreSQL refuses a statement that so much as *names* the column, even inside an ``IS NULL`` test, so there is no read path to narrow later. .. note:: That absoluteness has a practical edge. A listing query cannot report "does this identity have a private half?" as ``private_wrapped IS NOT NULL``: that is refused for every role but ``pepsi-crypto``, and PostgreSQL's error, ``permission denied for table crypto_identity``, names no column. The listing derives the flag from ``wrap_key_id`` instead, which a table ``CHECK`` keeps exactly in step. **By wrapping.** What is stored is not a private key but an AEAD ciphertext, bound to that row's ``(address, protocol, purpose)`` as associated data under a key-encryption key that lives **outside** the database, in a ``secrets.d/*.secret`` fragment readable only by ``pepsi-crypto`` — not by ``pepsi-owner``, so the one other role that can read the column still holds nothing it can open. Moving a blob between rows fails to authenticate. .. _sec-otp: A second factor for key management ================================== .. rubric:: Upholds :ref:`tm-adv-password` (a stolen mail password cannot register a key for a user who enrolled a second factor). A user may enrol a TOTP secret (RFC 6238: HMAC-SHA-1, six digits, 30-second steps — what every authenticator app implements) for their address; from then on every change to their keys that *they* can make needs a current code: * **By e-mail**, every key command but ``status`` — ``generate``, ``register``, ``publish``, ``retire`` and ``otp enroll``/``otp remove`` themselves — takes the code as a six-digit word in the ``Subject:``. An ordinary message's ``Autocrypt:`` header or attached key no longer registers anything for that address. * **In the web console and the API**, an ``own:
`` principal's ``generate-identity``, ``register-client-key`` and ``revoke-identity`` requests carry an ``otp`` field. ``pepsi-httpd`` cannot check it — it holds no grant on the table at all — and passes it through; ``pepsi-setup apply`` checks it before acting, for a task admitted on ``own:
`` alone. An operator holding ``keys:write`` is not asked. * **An unprivileged** ``pepsi-keys`` (a setuid install) needs ``--otp CODE`` for every identity change and private-key export; ownership is checked first, so one user can never count failures against another's enrolment. **Custody.** The secret is wrapped exactly like a private key — AES-256-GCM under the key store's KEK, bound to ``(address, totp, second-factor)`` — and the ``otp_key`` table is granted to the ``pepsi-crypto`` role alone: ``pepsi-setup`` revokes it from every other role, the pipeline's ``pepsi`` included, and verifies that against PostgreSQL. A stage that parses hostile mail can therefore neither read a secret nor reset a counter nor delete an enrolment (which would switch the protection off). ``pepsi-keys wrap rotate`` re-wraps the secrets with the private keys. **Attempts.** The code is checked by the process that can unwrap the secret; the *decision* is the ``otp_attempt`` SQL function's, under the row lock. A code matching a time step later than the last accepted one is accepted and resets the failure counter to zero; the same code again, or an older one, is a replay and counts as a failure, as does a wrong code. At **ten consecutive failures** the enrolment is locked and every attempt is refused unchecked. A command with no code at all is refused without being counted. Each enrolment has a never-reused serial, so an attempt checked against a secret that was replaced meanwhile is discarded rather than recorded against the new one. **The operator** clears a lock by removing the enrolment (``pepsi-keys otp reset``, or ``DELETE /api/v1/otp/{address}``, which queues a ``reset-otp`` task that only ``keys:write`` may ask for) or replacing it (``pepsi-keys otp enroll`` as the operator needs no code); a new enrolment always starts at zero failures. **What it does not do.** It does not protect the first enrolment (there is no code to ask for yet), nor an enrolment reply the user left in the mailbox. It does not gate the web console's synchronous flag changes (WKD publication, a key-server upload request, the primary MTA key), which ``pepsi-httpd`` performs itself. And it adds nothing against the operator, who can reset it. .. _sec-db: An attacker who has the database ================================ .. rubric:: Upholds :ref:`tm-adv-db`. Assume a full dump — a stolen backup, a replica, SQL injection into a read-only path. **They get:** * every message currently in the queue, in the clear: ``pepsi.workqueue`` holds the ``headers`` and ``pepsi.workqueue_body`` the body, unencrypted. Pepsi is a *pipeline*; a message in flight has to be readable by the stage that will process it. Delivered messages are deleted, so this is the backlog, not the history. * the full correspondence graph: senders, recipients, timestamps, per-address settings, whitelists, the ``origin_nonce``, ``auto_pay_spend``, ``payment_request_reply`` and ``dns_address`` tables. * every **public** key and certificate, and all key metadata. * the ARC/DKIM/DMARC verdicts and TLS session outcomes. **They do not get:** * **any private key**, nor any second-factor secret. The wrapped blobs are AEAD ciphertext and the key-encryption key is not in the database. * **any secure-link message body.** ``secure_message.ciphertext`` is AES-256-GCM under a key derived by Argon2id from the recipient's PIN, a per-message salt *and* a server-side pepper held in the configuration, not the database. The PIN itself is **never stored in any form** — not even a hash. A database thief must brute-force a PIN they cannot verify offline without also having the pepper. * **any usable credential.** Admin passwords are Argon2id hashes; API tokens are stored as digests, so a token cannot be recovered from a backup and is shown exactly once. * **the DKIM signing keys**, which live in the filesystem, and the SRS and Pepsi-Origin secrets, which live in ``secrets.d``. .. warning:: The queue is the exposure. An operator who considers mail-in-flight confidentiality important should encrypt the database at rest and keep backups accordingly — Pepsi cannot encrypt the queue to itself, because every stage would need the key, which is the same as not having one. .. _sec-config: An attacker who has the configuration file ========================================== .. rubric:: Upholds :ref:`tm-adv-secret`. ``/etc/pepsi/pepsi.conf`` is mode **0644** on purpose: unprivileged stages read it. It therefore must contain **no secrets**, and does not — each secret lives in a ``secrets.d/.secret`` fragment pulled in by an ``@inline-secret@`` directive and owned by the one account that needs it. **Read access to the config file** yields the topology: hostnames, served domains, the pipeline graph, smarthost addresses, file paths, which features are on. That is reconnaissance, not compromise. **Read access to a ``secrets.d`` fragment** yields exactly what its one reader holds, and no more — the SRS key, or the smarthost password, or the key-encryption key, or the secure-link pepper, or the list unsubscribe key. That is why they are split: there is no single file whose disclosure is total. **Write access to the config file** is equivalent to root. It names the programs each stage runs. Treat it as a root-owned file, because it is one. .. warning:: Two traps around ``@inline-secret@``, both of which look like a working configuration: * the directive **ends its section** — anything written after it lands in no section at all; * an **unreadable** fragment only *warns*. The option then reads as unset, and a feature silently runs without its secret. Verify by ``stat``-ing the fragment as the reader account, not by reading the log. .. _sec-worker: An attacker who compromises a stage worker ========================================== .. rubric:: Upholds :ref:`tm-adv-component`, including its signing residual. Say a parser bug gives code execution inside a stage — the realistic case, since stages parse hostile input for a living. **A ``pepsi``-user stage** (most of them) gets the queue: read, modify, redirect or delete any message in flight, and read the whitelist and settings. It does **not** get the wrapped end-to-end private keys, and it cannot become another account. It **does**, however, get the ability to have mail *signed*, and this is worth being blunt about because it is the largest residual in the design: * ``pepsi-stage-dkim-sign`` runs as ``pepsi`` and reads the domain DKIM keys through group ``pepsi``, so any ``pepsi`` process can sign as every served domain — directly, or by editing a queued row before the signing stage sees it. * ``pepsi-stage-encrypt`` signs as the row's ``From:`` author. Ingress enforces ``USERNAME_MAP`` at submission, but the queue row is writable by every ``pepsi`` stage afterwards and nothing downstream can tell whether ingress made that decision. A compromised ``pepsi`` stage can therefore rewrite a queued submission's ``From:`` (or inject a message with ``workqueue_inject``) and have it signed with **any local user's end-to-end key**. So the integrity of a Pepsi-made signature — DKIM or per-user OpenPGP/S/MIME — is only as strong as the whole ``pepsi`` pipeline role, not as strong as ``pepsi-crypto``. What the split *does* bound is **key disclosure**: a compromised language detector can get messages signed while it is running, but it cannot carry a private key away, so the damage ends when the stage is fixed and does not require re-keying. See :doc:`threat-model` for where this sits among the other guarantees. **A ``pepsi-crypto`` stage** gets everything that account has: it can decrypt inbound mail addressed to the deployment and sign as any local identity. This is the highest-value target in the system, which is why the two crypto stages are small, do no network I/O at all (discovery is a separate program), and are the subject of the fuzzing in :doc:`testing`. **``pepsi-ingress``** sees every message as it arrives and decides who submitted it, so it can admit mail as any local user. Its database role can add to the queue and read nothing of it, so what is already queued, the settings, the whitelists and the keys stay out of its reach. **``pepsi-httpd``** gets the web surface and its database grants — envelopes and routing of the queue, never a message's headers or body. It notably *cannot* write configuration or run privileged setup steps: the browser wizard does not act, it inserts a row into ``pepsi.setup_task`` describing what should be true, and the separately-started, root-running ``pepsi-setup apply`` decides whether to do it. A compromised web tier can therefore *request* a reconfiguration but not perform one — and the applier validates each task rather than trusting its description. That last defence can be made absolute, which is the right posture on any deployment administered from a terminal. The applier's two systemd units are their own Debian package, ``pepsi-httpd-admin`` (``pepsi`` only ``Recommends`` it, so ``apt remove pepsi-httpd-admin`` succeeds and leaves the mail server running); a source install gets the same lever from ``make install INSTALL_ADMIN_UNITS=no``. With them gone **nothing drains** ``pepsi.setup_task``, so an insert into it is inert whatever happens to ``pepsi-httpd`` — the loop is broken at the mechanism rather than at the interface, and no code path in the web tier can start a root process. The console detects this, fails closed if it cannot tell, and serves its setup pages read-only. ``pepsi-setup`` stays in the base package, so administering the host by hand is unaffected. See :ref:`packages-hardening`. Specific defences ================= The submission identity ----------------------- .. rubric:: Upholds :ref:`tm-adv-local` and :ref:`tm-adv-mailonly` (boundary B2): a submitter sends only as the identities the operator mapped to them. Ingress decides who submitted a message by one of four means — SASL through the configured backend, a TLS client certificate, an address in ``MYNETWORKS``, or the ``SO_PEERCRED`` credentials of the local submission socket, which is how ``sendmail`` authenticates without a password and why that socket may be world-writable. Whichever it is, ``USERNAME_MAP`` then decides which ``From:`` and envelope addresses that identity may use (RFC 6409 §6/§8.1), and an identity mapped to nothing may authenticate but not send. Speaking *for* an address — registering a key for it, or operating its keys by e-mail — needs more than being permitted to send as it: the ``From:`` must be exactly one mailbox named by a concrete (wildcard-free) entry of the map (``state.origin.from_bound``). A ``*@example.org`` grant lets an account send as every colleague, never plant a key for them. The decision is recorded in the message's ``state`` (``local_origin``, the authenticated identity) and believed by every later stage — which is exactly why :ref:`tm-adv-component` treats a compromised ``pepsi`` stage as able to forge it. Re-sealing decrypted mail for storage ------------------------------------- .. rubric:: Upholds :ref:`tm-client` (mail the gateway opened is filed readable only by the user's own client) and :ref:`tm-matrix` (the *Filed mail* column). ``pepsi-stage-decrypt`` has to produce plaintext: the spam, language, whitelist and key-learning stages after it read the content. ``pepsi-stage-reencrypt`` runs once they have, immediately before local delivery, and seals the message to the recipient's newest active ``custody = client`` key. The mechanism rests on four choices: * **It needs no privilege.** Sealing to a *public* key uses no private material, so the stage runs as ``pepsi``, is folded into the multi-call binary, and reads only the public columns of ``crypto_identity`` under the ordinary pipeline grant. It never names ``private_wrapped``, and gaining the ``pepsi-crypto`` identity would give it nothing it uses. * **It signs nothing.** A gateway signature would only say "the gateway sealed this", indistinguishable in a client from authorship, and would need the setuid identity. The sender-signature verdict is the one decrypt recorded, in ``state.crypto.in`` and in ``X-Pepsi-Crypto`` — which stays on the outer block and which decrypt strips from every inbound message before writing its own, so it is still ours when the client reads it. * **It acts only on what decrypt opened** (``state.crypto.in`` says encrypted and decrypted), and records ``state.crypto.store``, which also stops it from ever sealing a row twice. Mail that arrived in cleartext is not sealed. * **A refusal returns nothing that was encrypted.** Under ``ON_NO_CLIENT_KEY = bounce`` the DSN's header block is cut to the addressing fields, the body is emptied and ``RET=HDRS`` forced, since the sender encrypted the message and a DSN travels in clear. It bounces rather than defers, because deferring would keep the plaintext queued longer. The residual is the one :ref:`tm-adv-component` already states: a compromised ``pepsi`` stage — the re-encryption stage included — sees the plaintext while it is queued and can route a message past the stage, so the guarantee is about copies *already filed*, which nothing on the server can open. What an S/MIME trust anchor vouches for --------------------------------------- .. rubric:: Upholds :ref:`tm-indicators`, against :ref:`tm-adv-remote`: a CA the operator trusts for one namespace cannot make ``valid`` appear for another. An organisation that exchanges signed mail with a partner often trusts the partner's CA only for the partner's own domain — by cross-signing it with ``nameConstraints``, or by anchoring a partner sub-CA that carries them. If the bound were ignored, that partner CA could issue a certificate for one of *this* deployment's users and the decrypt stage would call the signature ``valid``. ``pepsi-crypto::trust`` therefore runs RFC 5280 §6.1's constraint state down every chain it builds: * **Name constraints** apply to the subject DN, every subject alternative name, and every ``emailAddress`` attribute in the subject DN — the last always, because that attribute is what a message is bound to when the SAN names no mailbox. Permitted subtrees intersect down the chain (two CAs permitting disjoint domains leave *no* mailbox permitted, not every one); excluded subtrees accumulate. A sub-CA cannot widen what its parent delegated. * **The anchor's own constraints bind it** (RFC 5937), so anchoring the partner's sub-CA directly is as safe as cross-signing it. * **Policy constraints** (``requireExplicitPolicy``, ``inhibitPolicyMapping``, ``inhibitAnyPolicy``) are evaluated over the full policy tree, in a deduplicated form whose size is polynomial in the policies a certificate names, and capped (64 policies or mappings per certificate) — the tree that grows exponentially in the naive algorithm is the OpenSSL CVE-2023-0464 denial of service. Name checking is capped at 2\ :sup:`20` comparisons per certificate for the same reason. * **What cannot be honoured is refused.** An unprocessable constraint — a subtree form Pepsi does not match, when a name of that form appears below it; a ``minimum`` or ``maximum``; an undecodable extension, critical or not — makes the chain ``policy-rejected`` rather than valid without the bound. Where Pepsi departs from RFC 5280 or from OpenSSL it is always in that direction: a wildcard ``dNSName`` that can stand for an excluded host is refused, and a host-less URI under a URI constraint is refused. The NIST PKITS conformance suite (sections 4.6 and 4.8–4.13, every input setting it lists) runs in ``cargo test``, alongside a corpus of our own for the forms PKITS does not cover (IP addresses, unimplemented forms, a name-constrained anchor). Header forgery -------------- .. rubric:: Upholds :ref:`tm-indicators`, against :ref:`tm-adv-remote`. ``X-Pepsi-Crypto`` tells a user's client what Pepsi concluded about a message's cryptography, so a forged one is a lie told with our authority. Inbound mail has the header **stripped before** the decrypt stage can add its own, and outbound request headers are stripped by ``STRIP_REQUEST_HEADERS``. A remote sender cannot make one appear. The same reasoning covers ``Authentication-Results`` and ``Pepsi-Origin``: a header that means "the local system verified this" is only meaningful if the local system is the only thing that can write it. Unknown recipients and backscatter ---------------------------------- .. rubric:: Upholds :ref:`tm-assets` (sending identity and reputation), against :ref:`tm-adv-remote`. A bounce goes to whatever envelope sender a message claims, and spam forges it. An MTA that accepts mail for addresses it does not have and bounces it later therefore mails strangers on a spammer's behalf. That is *backscatter*, and blocklists act on it quickly. Pepsi closes this in three layers, each a net under the one before: * **Refuse at RCPT.** ``pepsi-ingress`` answers ``550 5.1.1`` for an address at a served domain that nothing on the host vouches for: a local account, an alias, a list, a service address, or the MDA or backend asked with ``RCPT TO``. The sources are derived from the stage graph (``pepsi-recipients``). Every uncertainty accepts, so an outage never turns into refused mail. * **Report only to a sender that proved itself.** ``pepsi-stage-bounce`` sends a DSN for an inbound message only when SPF passed for the envelope sender's domain or an aligned DKIM signature did (``BOUNCE_UNAUTHENTICATED = drop``, the default). Anything else is dropped and logged. * **Refuse a loop.** A message that has already been through this host's ingress ``MAX_SELF_HOPS`` times is refused as a routing loop. An address on our own domain relayed to our own MX therefore costs one bounce, not a trip round the pipeline for every hop the relay stages allow. Answering ``RCPT`` truthfully tells a sender which addresses exist, as every MTA that rejects unknown users does. ``MAX_INVALID_RECIPIENTS`` closes a session after ten wrong guesses, so harvesting costs connections, which the connection limits meter. Probes of the MDA or backend are cached and capped at 16 in flight (beyond that the answer is "unsure", which accepts). A flood of guesses therefore cannot be turned against the mail store either. ``VRFY`` still answers ``252`` for everything. The recipient check needs one grant that ``pepsi-ingress`` did not have: the name and host columns of ``mailing_list``, which together are a list's public address and nothing more. A correspondent's key changing ------------------------------ .. rubric:: Upholds :ref:`tm-peer-keys`, against :ref:`tm-adv-remote`: a harvested key may be replaced by a newer one, but never silently. Inside the Autocrypt tier (the ``inbound`` and ``gossip`` rungs) the conflict rule is Autocrypt Level 1's newest-wins, keyed on the ``From:`` header a sender writes and on a message that inbound authentication accepted fail-open. So a message that only *claims* to come from a correspondent can replace the key we encrypt to them with. The mechanisms that bound this: * **The replacement is recorded by the statement that makes it.** ``crypto_peer_key_upsert`` writes a ``key.peer.rotate`` row to ``pepsi.event_log`` (the append-only audit log no mail-handling role can ``UPDATE`` or ``DELETE``) in the same transaction as the swap, naming the address, both fingerprint sets and both effective dates. No caller can rotate a key without the row, and a compromised stage that wanted to hide one would have to be the kind of adversary :ref:`tm-adv-component` already concedes the queue to. * **The people about to reply are told.** ``pepsi-stage-autocrypt-learn`` mails each local envelope recipient of the rotating message (``RESPONSE_STAGE``, the ``key-rotation`` template) with both fingerprints, so a forged rotation is something the targeted user can notice before sending anything confidential. Only recipients at served domains are mailed, so the notice cannot be turned into backscatter, and at most one per recipient per message. * **An operator can refuse rotation over a live key.** ``ACCEPT_ROTATION = expired`` accepts the newer key only when every stored one is recorded as revoked or expired, or is past the expiry its own material states (stored when the key was learnt, and for OpenPGP only ever moved later, so a replayed older copy of the certificate cannot backdate it); otherwise the offer is held, audited as ``key.peer.rotate.held`` and reported to the recipients. It trades a forged rotation for a genuine one refused until an operator intervenes. * **The tier stays confined.** None of this widens what may be replaced: nothing a domain, a directory, a key server or an operator published, and no pinned row, can be touched by a message, and an address we hold an identity for is pre-empted by our own key regardless. A served address we hold *no* identity for is not pre-empted, and ``pepsi-keys peer unclaimed`` lists every such address with a harvested or gossiped key (:ref:`km-unclaimed`). * **Promotion is not rotation.** When a gossiped key's owner advertises the same key themselves, ``crypto_peer_key_upsert`` moves the row from the ``gossip`` to the ``inbound`` rung and writes a ``key.peer.promote`` row (never ``key.peer.rotate``, and no notice, since the key did not change). It is decided under the same per-address lock, keyed on the source *names*, and maps to the same ``TrustSource::Autocrypt`` / ``Untrusted`` verdict as any harvested key (:ref:`km-gossip-promotion`). ``pepsi-status`` lists the last 30 days of both kinds of row. Key-discovery SSRF ------------------ .. rubric:: Upholds :ref:`tm-adv-remote` and :ref:`tm-peer-keys`: an address somebody writes to cannot aim the server at its own network. Key discovery fetches URLs derived from an e-mail address — WKD, and HTTP key servers — which is a server-side request forgery primitive by construction. ``pepsi-keydisc`` resolves the hostname itself and refuses **every** candidate address in loopback, RFC 1918, RFC 6598 CGNAT, link-local or unique-local space, and re-applies the check on each redirect (the redirect is the hole an address-check-then-connect design leaves open). The check is disabled only by an explicit test seam, because a stub HTTPS server necessarily listens on loopback. The secure-link portal ---------------------- .. rubric:: Upholds :ref:`tm-adv-db` and the stored-message claim of :doc:`secure-link`. The portal renders attacker-influenced content (a subject line, a sender name) into HTML, and gates it on a PIN. The design decisions that matter: * the **PIN is never stored**, in any form. It is an input to the Argon2id derivation of the content key, so an attacker with the database cannot verify a guess without also having the server pepper; * the **link and the PIN travel separately** — the link to the recipient, the PIN to the sender to pass on by another channel. Compromising one mailbox yields neither the message nor a way to get it; * wrong PINs are counted and lock the message out for a configured interval, so the online guessing rate is bounded; * a reply may be addressed **only** to the original sender, recorded in the row. The portal is not a mail-sending surface. The administrative API ---------------------- .. rubric:: Upholds :ref:`tm-adv-delegated` and boundaries B6/B7. Three authentication mechanisms, one identity model: ``SO_PEERCRED`` over a UNIX socket (unforgeable, and how first-run bootstrap works with no accounts yet), bearer tokens for automation (digest-stored, constant-time compared), and password plus session cookie for browsers (Argon2id, ``HttpOnly`` and ``SameSite=Strict``, CSRF token on every mutating request). Administrative routes mount **only** on listeners flagged ``ADMIN = yes``, so a relay install that never enables such a listener does not expose the surface at all. Responses never contain secrets: ``GET /api/v1/config`` masks credential-bearing options rather than omitting them, so an operator can see that a password is set without the endpoint becoming a way to read one back. Switching feature telemetry on from the console ----------------------------------------------- .. rubric:: Upholds :ref:`tm-assets` (privacy toward third parties: nothing the operator did not opt into) and :ref:`tm-adv-component` (``pepsi-httpd``). ``[pepsi] SHARE_TELEMETRY`` lives in the configuration file, so the console changes it only the way it changes anything there: a ``write-config`` request that root's applier performs. What makes that safe to offer is a precondition the operator sets up by hand and the console can only *observe*: ``pepsi-telemetry-client`` must be running (dormant while telemetry is off) and ``[pepsi] SYSTEM_ID`` must be present. The daemon reports both in the one-row ``pepsi.telemetry_client`` table, which ``pepsi-httpd`` may read and not write — so a console cannot forge "ready" — and the row carries a boolean for the identifier, never its value. Without it the question is disabled with directions and an API request to switch on is refused (``409 telemetry_client_not_ready``); switching *off* is never refused. That gate is a statement of usefulness, not the boundary. The boundary is the daemon: it believes only the file (through the single reader of the switch), re-reading it on the ``telemetry_changed`` notification, on ``SIGHUP`` and every minute; a notification cannot turn it on. A compromised ``pepsi-httpd`` holding ``CONFIG_DB`` could already request any configuration, this switch included, and that is unchanged; what it cannot do is make a deployment whose operator never started the daemon submit anything. Nothing in this path mints an identifier the installation did not already have: the applier carries an existing ``SYSTEM_ID`` across a console save instead, and ``pepsi-setup`` generates one only on a *yes*. Switching off discards what the daemon had gathered and not yet sent. Resource limits --------------- .. rubric:: Upholds :ref:`tm-availability`: a stranger cannot make one message, one connection or one address cost the server without bound. The limits themselves are tabulated in :ref:`tm-availability`; what matters here is where each is enforced and why it cannot be stepped around. * **Connections are bounded per source, not only in total.** ``pepsi-ingress`` counts open connections per source range (``MAX_CONNECTIONS_PER_IP``) and, for the UNIX-socket peers that are exempt from its rate limit and never evicted, for all of them together (``MAX_LOCAL_CONNECTIONS``); ``pepsi-httpd`` does the same per client address. The check happens before a connection takes a slot of the global ceiling, so a source at its share cannot also queue for more. * **Guessing and flooding are counted across sessions.** A per-range token bucket (``AUTH_FAILURES_PER_HOUR``) is spent by every failed ``AUTH``, and once empty the next ``AUTH`` is refused *before* the SASL backend is asked; a per-uid bucket keyed on the kernel-reported ``SO_PEERCRED`` uid (``LOCAL_MESSAGES_PER_MINUTE``) meters local submission, and a session may start only ``MAX_MESSAGES_PER_SESSION`` transactions before it must reconnect into the rate limit. * **One sender's burst does not delay everybody else.** ``pepsi-dispatch`` claims each stage's work by least virtual runtime per sender rather than oldest first (``FAIR_SCHEDULING``, on by default). The sender is the generated ``workqueue.sender_key``: the authenticated account for a submission, otherwise the connecting address — an IPv6 peer by its /64, since a single host is routinely given one — and never the envelope domain of mail from outside, which the sender chooses per message. Worker time is charged to the sender that caused it, a newcomer starts level with the least-served sender so silence banks no credit, and a reserved oldest-first share (``FAIR_FIFO_PERCENT``) bounds any one message's wait. The accounting is in the dispatcher's memory only; it grants nothing and changes no grant. * **Inbound verification has a work bound that is not a way around DMARC.** At most ``MAX_DKIM_SIGNATURES`` signatures are verified, those that can align with the ``From:`` domain first, and SPF and each signature run concurrently under ``VERIFY_TIMEOUT``, with DMARC under a second deadline. A check that runs out of time becomes ``temperror`` *for that check*: DMARC consults a temporary error only on an identifier that aligns with the ``From:`` domain, and the DNS DMARC itself reads is that domain's own, so a forger who slows their own DNS down gains nothing but their own ``temperror``. * **Key material from strangers has a size ceiling.** Every OpenPGP certificate that arrives from outside — WKD, a key server, DANE, a harvested or gossiped header, an attached key, a pasted one — passes through one parser that refuses an RSA key above 8192 bits, and the verifier drops such a certificate from its trust store as well. S/MIME has the same bound at 4096 bits, from the ``rsa`` crate. * **Retention does not depend on the web tier.** The audit and mail logs are pruned by ``pepsi-httpd prune`` from ``pepsi-log-prune.timer`` — as the ``pepsi-httpd`` account, the one service role holding ``DELETE`` on them, which keeps the logs append-only for everything that handles mail — and TLS-report counters and expired MX cache rows by their own timers, all named by ``pepsi.target``. * **The queue has a ceiling, and the check cannot be starved or bypassed by splitting.** ``pepsi-ingress`` refuses with ``452 4.3.1`` at ``RCPT``, ``DATA``/``BDAT`` and after the end of data while the queue is at ``[pepsi] MAX_QUEUE_ROWS``/``MAX_QUEUE_BYTES`` or the database's filesystem below ``MIN_FREE_SPACE``. The measurement (``pepsi.queue_usage()``) reads only two generated size columns, so the least-privileged ingress role gains no access to queued mail for it, and it is cached per process (``QUEUE_CHECK_INTERVAL``), so a flood of ``RCPT`` commands costs no queries. Bytes are counted per *stored* body: bodies live in ``pepsi.workqueue_body``, shared by every row a split, clone or list fan-out makes, so one large message to many recipients can no longer multiply its body across the queue — the amplification the per-recipient split used to offer. The one amplifier left under a sender's control, a post to a large list, holds its fan-out back (``pepsi-stage-list-post``, ``QUEUE_THROTTLE_DELAY``) while the queue is full. Deleting the last row that references a body deletes it (a trigger), and the foreign key means no role — the web tier's ``DELETE`` on ``workqueue_body`` included — can delete a body a queued message still carries. * **Mailing lists bound what one sender can make them do.** A sender's accepted posts to a list are counted per hour (``SENDER_POSTS_PER_HOUR``) in one SQL statement that decides and records, so parallel workers cannot both take the last slot; the excess is held rather than multiplied by the membership. The moderation queue is capped per list (``MAX_HELD_MESSAGES``) by discarding the *new* post, so a flood cannot flush out the posts a moderator has yet to see. * **Automatic spending has an aggregate, and it cannot be raced.** ``pepsi-stage-auto-pay`` books a demand against two budgets in the one ``auto_pay_try_spend`` call that also records it: the message's (``MAX_TOTAL``, on the ``origin_nonce`` row, locked ``FOR UPDATE``) and the paying wallet's rolling 24 hours (``MAX_TOTAL_PER_DAY``, the ``auto_pay_spend`` ledger, under a transaction-scoped advisory lock keyed on the wallet — demands for different messages share no row, so without it each would read the same sum). A refusal by either writes nothing to the other; a payment that fails to settle is credited back to both by ``auto_pay_refund``, which takes only the first lock, so the two cannot deadlock. The wallet key is the one the helper pays from (``uid:``, or the shared account and wallet database), so two senders cannot share a day unless they share a wallet. * **A payment request answers a sender once per window.** The request is an automatic reply to an unverified envelope sender, so besides RFC 3834 suppression ``pepsi-stage-anti-spam`` claims it through ``payment_request_should_send`` — ``INSERT … ON CONFLICT DO UPDATE … WHERE`` deciding and recording in one statement, as ``vacation_should_reply`` does — keyed on the protected mailbox and the sender, with ``BLOCK_RESPONSE_SUPPRESS`` as the window. A failure to claim withholds the reply; the message is held either way. What Pepsi does not defend against ================================== The list is kept in one place, :ref:`tm-out-of-scope`, together with the residuals each adversary section states: root and a malicious operator, traffic analysis and metadata, plaintext in the queue, signatures against a compromised pipeline, correspondents' own compromise, and a network attacker who can also lie in DNS. .. _reporting-a-vulnerability: Reporting a vulnerability ========================= Please report security issues privately by e-mail to Christian Grothoff at grothoff@gnunet.org, encrypted to the OpenPGP key published at https://grothoff.org/christian/grothoff.asc, whose fingerprint is ``D842 3BCB 326C 7907 0339 29C7 939E 6BE1 E29F C3CC``. Do not use the public bug tracker for a vulnerability, and allow time for a fix before disclosure. ``SECURITY.md`` in the source tree states the full policy. Ordinary bugs go to the public bug tracker at https://bugs.gnunet.org/view_all_bug_page.php?project_id=35. See also ======== :doc:`key-management` — identities, custody, discovery and the trust ladder. :doc:`secure-link` — the portal in full. :doc:`admin-api` — the API surface and its scopes. :doc:`interoperability` — what other clients can actually read, which bounds what protection is achievable in practice. :doc:`testing` — the fuzzing and adversarial corpora behind the parser claims.