.. This file is part of PEPSI. Copyright (C) 2026 GNUnet e.V. PEPSI is free software; you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation; either version 3, or (at your option) any later version. PEPSI is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU Affero General Public License for more details. .. _operations: =============== Operating Pepsi =============== This chapter is about the weeks after installation: where the logs are, what to watch, and what to back up so that a lost disk costs you an afternoon rather than every user's keys. Upgrading is described in :ref:`upgrading`, and what to do when something goes wrong in :doc:`troubleshooting`. .. _operations-logs: Logs ==== Every Pepsi program logs to standard error. Under the shipped systemd units that is the journal, one unit per long-lived program: .. code-block:: console journalctl -u pepsi-ingress -u pepsi-dispatch -u pepsi-httpd --since today journalctl -u pepsi-dispatch -f # follow the pipeline live journalctl -u 'pepsi-*' -p warning # warnings and errors, every unit The stage programs have no units of their own: the dispatcher starts them as workers and they inherit its standard error, so **everything a stage logs appears under** ``pepsi-dispatch``. Each record names the program that wrote it. The timers (key refresh, quota reconciliation, log pruning, …) log under their own ``.service`` names; ``systemctl list-timers 'pepsi-*'`` lists them with their last and next run. Log levels ---------- The levels are ``error``, ``warn``, ``info`` (the default), ``debug`` and ``trace``. * ``[pepsi] LOG`` sets the level for **every** program at once, the stage workers included, which take no command-line flags from the dispatcher. Restart the services after changing it (the dispatcher restarts its workers itself). * ``-L LEVEL`` (before the subcommand, e.g. ``pepsi-queue -L debug list``) overrides it for one process, which is the usual way to look closer at one tool run by hand. * ``-v`` also lets through the database, HTTP and TLS libraries' own records, which are hidden otherwise. Useful for a connection problem; very noisy. * ``[pepsi] LOG_JSON = yes`` writes one JSON object per record, for a log aggregator. ``info`` records one line or a few per message and stage. ``warn`` keeps what needs attention and suits a busy host. ``debug`` and ``trace`` can include envelope addresses and protocol detail; do not leave them on. What is kept, and for how long ------------------------------ The journal's retention is systemd's (``journald.conf``). Inside the database Pepsi keeps: * the **audit log** (``pepsi.event_log``): who changed what, from the command-line tools, the API and the console, pruned daily by ``pepsi-log-prune.timer``; * **no per-message delivery log**, unless you turn on ``[pepsi] MAIL_LOG`` (see :manpage:`pepsi.conf(5)` and the privacy warning there). A delivered message is deleted from the queue, so without it the journal is the only record that a message passed through; * TLS reporting counters (``pepsi.tls_session``), pruned by ``pepsi-tlsrpt-prune.timer``. .. _operations-monitoring: Monitoring ========== Two things tell you whether mail is flowing: the queue, and the messages the pipeline has given up on. ``pepsi-status`` ---------------- :doc:`programs/pepsi-status` is a read-only summary of the queue, the stuck messages with the reason for each, cumulative per-stage counters, outbound TLS results, mailbox quotas and recent key changes. It is safe to run at any time, and ``--json`` gives the same report to a script: .. code-block:: console pepsi-status pepsi-status --json | jq .problem_total ``problem_total`` is the number of messages in ``failed`` or ``timeout``. The dispatcher never retries them; ``pepsi-failure-bouncer.timer`` bounces them to their senders once they have been failed for ``MIN_AGE`` (an hour; see :ref:`troubleshooting-failed`), so the count normally drops back to zero by itself. If it keeps rising, the same stage is failing message after message and the cause needs fixing. A host fault does not show here at first: its messages wait ``paused`` at the failing stage (``pepsi_pause_backlog``) with the error in ``state.last_error``, and are retried until the stage's ``MAX_LIFETIME``. ``/metrics`` ------------ ``pepsi-httpd`` serves Prometheus metrics at ``/metrics`` on a listener flagged ``ADMIN = yes`` (and nowhere else; it takes no credential, so flag a listener only your monitoring system can reach). The series: .. list-table:: :header-rows: 1 :widths: 40 12 48 * - Series - Type - Meaning * - ``pepsi_stage_active_messages{stage}`` - gauge - Messages being processed right now. * - ``pepsi_pause_backlog{stage}`` - gauge - Messages waiting for a retry: a remote server that deferred, a key being looked up, a payment awaited. * - ``pepsi_stage_messages_total{stage}`` - counter - Messages a stage has processed. * - ``pepsi_stage_duration_seconds_total{stage}`` - counter - Time spent in the stage; divided by the previous series, the average. * - ``pepsi_stage_timeouts_total{stage}`` - counter - Workers killed for exceeding the stage's ``MAX_RUNTIME``. * - ``pepsi_stage_crashes_total{stage}`` - counter - Workers that died while handling a message. * - ``pepsi_messages_processed_total`` - counter - Messages that left the pipeline. * - ``pepsi_stages_executed_total`` - counter - Stage passes, all stages together. The counters are written by the dispatcher every ``STATS_INTERVAL`` and when it goes idle, so they lag by up to that interval. They are totals since the schema was created, not per scrape; use ``rate()``. What to alert on ---------------- .. list-table:: :header-rows: 1 :widths: 40 60 * - Condition - Why * - ``pepsi-status --json`` reports ``problem_total`` above zero for longer than ``MIN_AGE`` plus ten minutes - A message has failed and has not been bounced: ``pepsi-failure-bouncer.timer`` is not running (``systemctl list-timers``), or the bounce itself failed. Messages that fail at the bounce stage are not bounced again and wait for you. * - ``pepsi_pause_backlog`` of a non-delivery stage growing, or the journal repeating "failed at stage … retrying" - A host fault at that stage (an unreadable key or template, a helper that cannot start, a refused statement). Its messages are held, not bounced, until ``MAX_LIFETIME``; ``state.last_error`` says why. * - The dispatcher logging "the worker refused to start" for a stage - The stage's configuration or a secret it reads no longer parses; its messages are held until it does. * - ``pepsi_pause_backlog`` of a relay stage growing for hours - The next hop, DNS or outbound port 25 is unreachable. Retries continue until ``MAX_LIFETIME``, then the senders get bounces. * - ``rate(pepsi_stage_crashes_total[15m])`` or ``rate(pepsi_stage_timeouts_total[15m])`` above zero - A stage program is failing; its messages end as ``failed`` or ``timeout``. * - ``rate(pepsi_messages_processed_total[1h])`` at zero when mail is expected - Nothing leaves the pipeline: the dispatcher is down or a stage is stuck. * - A ``pepsi-*`` unit in ``failed`` state (``systemctl --failed``) - The long-running units retry forever, so a failed one is usually a timer or a one-shot job. * - Journal records at ``error`` - Anything logged as an error needs looking at. * - The TLS certificates' expiry - ``certbot`` renews them; a renewal that fails is logged by certbot, not by Pepsi. * - Free space where PostgreSQL keeps its data - Below ``[pepsi] MIN_FREE_SPACE`` (default 1 GiB), ``pepsi-ingress`` refuses new mail with a temporary error. .. _backup-restore: Backup and restore ================== .. warning:: **A database dump alone is not a backup of Pepsi.** The private keys in the database are encrypted under a key-encryption key that is kept, on purpose, *outside* the database, in ``/etc/pepsi/secrets.d/pepsi-crypto.secret``. Lose that file and every user's stored private keys are lost with it: nothing can open them again, and ``pepsi-setup`` will not generate a replacement while encrypted keys exist. What to back up --------------- .. list-table:: :header-rows: 1 :widths: 32 68 * - What - What is lost without it * - ``/etc/pepsi/secrets.d/``, above all ``pepsi-crypto.secret`` - ``pepsi-crypto.secret``: every stored private key (above). The other fragments: the SRS key (bounces to mail forwarded before the loss are refused), the secure-link pepper (secure messages not yet collected cannot be opened), the proof-of-origin key (payment demands for mail sent before the loss are no longer recognised as ours), and the passwords and tokens you entered. * - The PostgreSQL database (``pepsi`` by default) - The queue, per-user settings, whitelists, mailing lists and their archive, the key store, quotas, administrator accounts and API tokens, the audit log. * - ``/etc/pepsi/`` (the rest) - The configuration. It can be written again with the wizard, but not quickly. * - ``/var/pepsi/keys/`` - The DKIM and ARC signing keys. New ones can be generated, but each has to be published in DNS again, and mail signed before that fails DKIM. * - ``/var/lib/pepsi-wallets/`` - With ``pepsi-stage-auto-pay`` in shared wallet mode: the wallets and the money in them. In ``local-user`` mode the wallets are in the users' home directories. * - ``/etc/letsencrypt/`` - Nothing permanent; certbot obtains new certificates. Keep it anyway if you publish DANE (TLSA) records for the current key. * - ``/var/pepsi/tokens``, ``/var/pepsi/token-refresh``, ``/var/pepsi/tls`` - Only with an OAuth or client-certificate smarthost: the refresh token has to be authorised again. The users' mailboxes (``Maildir``, or the MDA's store behind LMTP) are not Pepsi's and belong in your ordinary backup of home directories. Keep the files and the database dump **together**, and from the same moment: a dump newer than ``secrets.d`` may hold keys wrapped under a key-encryption key the file backup does not have. Both hold secrets; store them encrypted and readable by root only. Taking a backup --------------- .. code-block:: console sudo -u pepsi-owner pg_dump --format=custom pepsi > pepsi-$(date +%F).dump tar -C / -czpf pepsi-files-$(date +%F).tar.gz \ etc/pepsi var/pepsi var/lib/pepsi-wallets Run as root, so that ``tar`` can read ``secrets.d``. ``pg_dump`` takes a consistent snapshot while the services run; the files change rarely (a new key, a configuration change), so taking them just before or after the dump is enough. The database name and its owner are those in ``[pepsi-postgres]``. The dump that ``pepsi-setup schema --backup-dir`` takes before an upgrade (see :ref:`upgrade-backup`) covers the database only, and only the ``pepsi`` schema; it does not replace this. Restoring --------- Restore on a host running **the same Pepsi release** the backup came from (upgrade afterwards, if you want to), in this order: 1. Install Pepsi. The Debian package creates the service accounts and database roles; on a source install create them as in :doc:`installation`. 2. Stop everything: ``systemctl stop pepsi.target``. 3. Restore the files, **secrets.d first**. ``tar`` keeps owners by name, so the accounts must exist before you unpack: .. code-block:: console tar -C / -xzpf pepsi-files-2026-10-01.tar.gz 4. Restore the database into an empty one owned by the schema owner: .. code-block:: console sudo -u postgres dropdb --if-exists pepsi sudo -u postgres createdb -O pepsi-owner pepsi sudo -u pepsi-owner pg_restore -d pepsi pepsi-2026-10-01.dump 5. Run ``pepsi-setup -c /etc/pepsi/pepsi.conf run``. It re-applies the role grants, writes again what lives outside both backups (the systemd drop-ins that hand the TLS certificates to the services) and checks the DNS records against the restored keys. It does not replace a key or secret that already exists. 6. Start the services: ``systemctl start pepsi.target``. If ``pepsi-setup`` says it refuses to generate a key-encryption key because identities exist, ``secrets.d/pepsi-crypto.secret`` was not restored, or is not the one that goes with this database. Do not work around it: find the right file.