7. Operating Pepsi

This chapter is about the weeks after installation: where the logs are, what to watch, and what to back up so that a lost disk costs you an afternoon rather than every user’s keys. Upgrading is described in Upgrading, and what to do when something goes wrong in Troubleshooting.

7.1. Logs

Every Pepsi program logs to standard error. Under the shipped systemd units that is the journal, one unit per long-lived program:

journalctl -u pepsi-ingress -u pepsi-dispatch -u pepsi-httpd --since today
journalctl -u pepsi-dispatch -f          # follow the pipeline live
journalctl -u 'pepsi-*' -p warning       # warnings and errors, every unit

The stage programs have no units of their own: the dispatcher starts them as workers and they inherit its standard error, so everything a stage logs appears under pepsi-dispatch. Each record names the program that wrote it. The timers (key refresh, quota reconciliation, log pruning, …) log under their own .service names; systemctl list-timers 'pepsi-*' lists them with their last and next run.

7.1.1. Log levels

The levels are error, warn, info (the default), debug and trace.

  • [pepsi] LOG sets the level for every program at once, the stage workers included, which take no command-line flags from the dispatcher. Restart the services after changing it (the dispatcher restarts its workers itself).

  • -L LEVEL (before the subcommand, e.g. pepsi-queue -L debug list) overrides it for one process, which is the usual way to look closer at one tool run by hand.

  • -v also lets through the database, HTTP and TLS libraries’ own records, which are hidden otherwise. Useful for a connection problem; very noisy.

  • [pepsi] LOG_JSON = yes writes one JSON object per record, for a log aggregator.

info records one line or a few per message and stage. warn keeps what needs attention and suits a busy host. debug and trace can include envelope addresses and protocol detail; do not leave them on.

7.1.2. What is kept, and for how long

The journal’s retention is systemd’s (journald.conf). Inside the database Pepsi keeps:

  • the audit log (pepsi.event_log): who changed what, from the command-line tools, the API and the console, pruned daily by pepsi-log-prune.timer;

  • no per-message delivery log, unless you turn on [pepsi] MAIL_LOG (see pepsi.conf(5) and the privacy warning there). A delivered message is deleted from the queue, so without it the journal is the only record that a message passed through;

  • TLS reporting counters (pepsi.tls_session), pruned by pepsi-tlsrpt-prune.timer.

7.2. Monitoring

Two things tell you whether mail is flowing: the queue, and the messages the pipeline has given up on.

7.2.1. pepsi-status

pepsi-status is a read-only summary of the queue, the stuck messages with the reason for each, cumulative per-stage counters, outbound TLS results, mailbox quotas and recent key changes. It is safe to run at any time, and --json gives the same report to a script:

pepsi-status
pepsi-status --json | jq .problem_total

problem_total is the number of messages in failed or timeout. The dispatcher never retries them; pepsi-failure-bouncer.timer bounces them to their senders once they have been failed for MIN_AGE (an hour; see Messages in failed or timeout), so the count normally drops back to zero by itself. If it keeps rising, the same stage is failing message after message and the cause needs fixing. A host fault does not show here at first: its messages wait paused at the failing stage (pepsi_pause_backlog) with the error in state.last_error, and are retried until the stage’s MAX_LIFETIME.

7.2.2. /metrics

pepsi-httpd serves Prometheus metrics at /metrics on a listener flagged ADMIN = yes (and nowhere else; it takes no credential, so flag a listener only your monitoring system can reach). The series:

Series

Type

Meaning

pepsi_stage_active_messages{stage}

gauge

Messages being processed right now.

pepsi_pause_backlog{stage}

gauge

Messages waiting for a retry: a remote server that deferred, a key being looked up, a payment awaited.

pepsi_stage_messages_total{stage}

counter

Messages a stage has processed.

pepsi_stage_duration_seconds_total{stage}

counter

Time spent in the stage; divided by the previous series, the average.

pepsi_stage_timeouts_total{stage}

counter

Workers killed for exceeding the stage’s MAX_RUNTIME.

pepsi_stage_crashes_total{stage}

counter

Workers that died while handling a message.

pepsi_messages_processed_total

counter

Messages that left the pipeline.

pepsi_stages_executed_total

counter

Stage passes, all stages together.

The counters are written by the dispatcher every STATS_INTERVAL and when it goes idle, so they lag by up to that interval. They are totals since the schema was created, not per scrape; use rate().

7.2.3. What to alert on

Condition

Why

pepsi-status --json reports problem_total above zero for longer than MIN_AGE plus ten minutes

A message has failed and has not been bounced: pepsi-failure-bouncer.timer is not running (systemctl list-timers), or the bounce itself failed. Messages that fail at the bounce stage are not bounced again and wait for you.

pepsi_pause_backlog of a non-delivery stage growing, or the journal repeating “failed at stage … retrying”

A host fault at that stage (an unreadable key or template, a helper that cannot start, a refused statement). Its messages are held, not bounced, until MAX_LIFETIME; state.last_error says why.

The dispatcher logging “the worker refused to start” for a stage

The stage’s configuration or a secret it reads no longer parses; its messages are held until it does.

pepsi_pause_backlog of a relay stage growing for hours

The next hop, DNS or outbound port 25 is unreachable. Retries continue until MAX_LIFETIME, then the senders get bounces.

rate(pepsi_stage_crashes_total[15m]) or rate(pepsi_stage_timeouts_total[15m]) above zero

A stage program is failing; its messages end as failed or timeout.

rate(pepsi_messages_processed_total[1h]) at zero when mail is expected

Nothing leaves the pipeline: the dispatcher is down or a stage is stuck.

A pepsi-* unit in failed state (systemctl --failed)

The long-running units retry forever, so a failed one is usually a timer or a one-shot job.

Journal records at error

Anything logged as an error needs looking at.

The TLS certificates’ expiry

certbot renews them; a renewal that fails is logged by certbot, not by Pepsi.

Free space where PostgreSQL keeps its data

Below [pepsi] MIN_FREE_SPACE (default 1 GiB), pepsi-ingress refuses new mail with a temporary error.

7.3. Backup and restore

Warning

A database dump alone is not a backup of Pepsi. The private keys in the database are encrypted under a key-encryption key that is kept, on purpose, outside the database, in /etc/pepsi/secrets.d/pepsi-crypto.secret. Lose that file and every user’s stored private keys are lost with it: nothing can open them again, and pepsi-setup will not generate a replacement while encrypted keys exist.

7.3.1. What to back up

What

What is lost without it

/etc/pepsi/secrets.d/, above all pepsi-crypto.secret

pepsi-crypto.secret: every stored private key (above). The other fragments: the SRS key (bounces to mail forwarded before the loss are refused), the secure-link pepper (secure messages not yet collected cannot be opened), the proof-of-origin key (payment demands for mail sent before the loss are no longer recognised as ours), and the passwords and tokens you entered.

The PostgreSQL database (pepsi by default)

The queue, per-user settings, whitelists, mailing lists and their archive, the key store, quotas, administrator accounts and API tokens, the audit log.

/etc/pepsi/ (the rest)

The configuration. It can be written again with the wizard, but not quickly.

/var/pepsi/keys/

The DKIM and ARC signing keys. New ones can be generated, but each has to be published in DNS again, and mail signed before that fails DKIM.

/var/lib/pepsi-wallets/

With pepsi-stage-auto-pay in shared wallet mode: the wallets and the money in them. In local-user mode the wallets are in the users’ home directories.

/etc/letsencrypt/

Nothing permanent; certbot obtains new certificates. Keep it anyway if you publish DANE (TLSA) records for the current key.

/var/pepsi/tokens, /var/pepsi/token-refresh, /var/pepsi/tls

Only with an OAuth or client-certificate smarthost: the refresh token has to be authorised again.

The users’ mailboxes (Maildir, or the MDA’s store behind LMTP) are not Pepsi’s and belong in your ordinary backup of home directories.

Keep the files and the database dump together, and from the same moment: a dump newer than secrets.d may hold keys wrapped under a key-encryption key the file backup does not have. Both hold secrets; store them encrypted and readable by root only.

7.3.2. Taking a backup

sudo -u pepsi-owner pg_dump --format=custom pepsi > pepsi-$(date +%F).dump
tar -C / -czpf pepsi-files-$(date +%F).tar.gz \
    etc/pepsi var/pepsi var/lib/pepsi-wallets

Run as root, so that tar can read secrets.d. pg_dump takes a consistent snapshot while the services run; the files change rarely (a new key, a configuration change), so taking them just before or after the dump is enough. The database name and its owner are those in [pepsi-postgres].

The dump that pepsi-setup schema --backup-dir takes before an upgrade (see Backing up before a schema upgrade) covers the database only, and only the pepsi schema; it does not replace this.

7.3.3. Restoring

Restore on a host running the same Pepsi release the backup came from (upgrade afterwards, if you want to), in this order:

  1. Install Pepsi. The Debian package creates the service accounts and database roles; on a source install create them as in Installation.

  2. Stop everything: systemctl stop pepsi.target.

  3. Restore the files, secrets.d first. tar keeps owners by name, so the accounts must exist before you unpack:

    tar -C / -xzpf pepsi-files-2026-10-01.tar.gz
    
  4. Restore the database into an empty one owned by the schema owner:

    sudo -u postgres dropdb --if-exists pepsi
    sudo -u postgres createdb -O pepsi-owner pepsi
    sudo -u pepsi-owner pg_restore -d pepsi pepsi-2026-10-01.dump
    
  5. Run pepsi-setup -c /etc/pepsi/pepsi.conf run. It re-applies the role grants, writes again what lives outside both backups (the systemd drop-ins that hand the TLS certificates to the services) and checks the DNS records against the restored keys. It does not replace a key or secret that already exists.

  6. Start the services: systemctl start pepsi.target.

If pepsi-setup says it refuses to generate a key-encryption key because identities exist, secrets.d/pepsi-crypto.secret was not restored, or is not the one that goes with this database. Do not work around it: find the right file.