7. Operating Pepsi¶
This chapter is about the weeks after installation: where the logs are, what to watch, and what to back up so that a lost disk costs you an afternoon rather than every user’s keys. Upgrading is described in Upgrading, and what to do when something goes wrong in Troubleshooting.
7.1. Logs¶
Every Pepsi program logs to standard error. Under the shipped systemd units that is the journal, one unit per long-lived program:
journalctl -u pepsi-ingress -u pepsi-dispatch -u pepsi-httpd --since today
journalctl -u pepsi-dispatch -f # follow the pipeline live
journalctl -u 'pepsi-*' -p warning # warnings and errors, every unit
The stage programs have no units of their own: the dispatcher starts them as
workers and they inherit its standard error, so everything a stage logs
appears under pepsi-dispatch. Each record names the program that wrote it.
The timers (key refresh, quota reconciliation, log pruning, …) log under their
own .service names; systemctl list-timers 'pepsi-*' lists them with
their last and next run.
7.1.1. Log levels¶
The levels are error, warn, info (the default), debug and
trace.
[pepsi] LOGsets the level for every program at once, the stage workers included, which take no command-line flags from the dispatcher. Restart the services after changing it (the dispatcher restarts its workers itself).-L LEVEL(before the subcommand, e.g.pepsi-queue -L debug list) overrides it for one process, which is the usual way to look closer at one tool run by hand.-valso lets through the database, HTTP and TLS libraries’ own records, which are hidden otherwise. Useful for a connection problem; very noisy.[pepsi] LOG_JSON = yeswrites one JSON object per record, for a log aggregator.
info records one line or a few per message and stage. warn keeps what
needs attention and suits a busy host. debug and trace can include
envelope addresses and protocol detail; do not leave them on.
7.1.2. What is kept, and for how long¶
The journal’s retention is systemd’s (journald.conf). Inside the database
Pepsi keeps:
the audit log (
pepsi.event_log): who changed what, from the command-line tools, the API and the console, pruned daily bypepsi-log-prune.timer;no per-message delivery log, unless you turn on
[pepsi] MAIL_LOG(see pepsi.conf(5) and the privacy warning there). A delivered message is deleted from the queue, so without it the journal is the only record that a message passed through;TLS reporting counters (
pepsi.tls_session), pruned bypepsi-tlsrpt-prune.timer.
7.2. Monitoring¶
Two things tell you whether mail is flowing: the queue, and the messages the pipeline has given up on.
7.2.1. pepsi-status¶
pepsi-status is a read-only summary of the queue, the stuck
messages with the reason for each, cumulative per-stage counters, outbound TLS
results, mailbox quotas and recent key changes. It is safe to run at any time,
and --json gives the same report to a script:
pepsi-status
pepsi-status --json | jq .problem_total
problem_total is the number of messages in failed or timeout. The
dispatcher never retries them; pepsi-failure-bouncer.timer bounces them to
their senders once they have been failed for MIN_AGE (an hour; see
Messages in failed or timeout), so the count normally drops back to zero by
itself. If it keeps rising, the same stage is failing message after message and
the cause needs fixing. A host fault does not show here at first: its messages
wait paused at the failing stage (pepsi_pause_backlog) with the error
in state.last_error, and are retried until the stage’s MAX_LIFETIME.
7.2.2. /metrics¶
pepsi-httpd serves Prometheus metrics at /metrics on a listener flagged
ADMIN = yes (and nowhere else; it takes no credential, so flag a listener
only your monitoring system can reach). The series:
Series |
Type |
Meaning |
|---|---|---|
|
gauge |
Messages being processed right now. |
|
gauge |
Messages waiting for a retry: a remote server that deferred, a key being looked up, a payment awaited. |
|
counter |
Messages a stage has processed. |
|
counter |
Time spent in the stage; divided by the previous series, the average. |
|
counter |
Workers killed for exceeding the stage’s |
|
counter |
Workers that died while handling a message. |
|
counter |
Messages that left the pipeline. |
|
counter |
Stage passes, all stages together. |
The counters are written by the dispatcher every STATS_INTERVAL and when it
goes idle, so they lag by up to that interval. They are totals since the schema
was created, not per scrape; use rate().
7.2.3. What to alert on¶
Condition |
Why |
|---|---|
|
A message has failed and has not been bounced:
|
|
A host fault at that stage (an unreadable key or template, a helper
that cannot start, a refused statement). Its messages are held, not
bounced, until |
The dispatcher logging “the worker refused to start” for a stage |
The stage’s configuration or a secret it reads no longer parses; its messages are held until it does. |
|
The next hop, DNS or outbound port 25 is unreachable. Retries continue
until |
|
A stage program is failing; its messages end as |
|
Nothing leaves the pipeline: the dispatcher is down or a stage is stuck. |
A |
The long-running units retry forever, so a failed one is usually a timer or a one-shot job. |
Journal records at |
Anything logged as an error needs looking at. |
The TLS certificates’ expiry |
|
Free space where PostgreSQL keeps its data |
Below |
7.3. Backup and restore¶
Warning
A database dump alone is not a backup of Pepsi. The private keys in the
database are encrypted under a key-encryption key that is kept, on purpose,
outside the database, in /etc/pepsi/secrets.d/pepsi-crypto.secret. Lose
that file and every user’s stored private keys are lost with it: nothing can
open them again, and pepsi-setup will not generate a replacement while
encrypted keys exist.
7.3.1. What to back up¶
What |
What is lost without it |
|---|---|
|
|
The PostgreSQL database ( |
The queue, per-user settings, whitelists, mailing lists and their archive, the key store, quotas, administrator accounts and API tokens, the audit log. |
|
The configuration. It can be written again with the wizard, but not quickly. |
|
The DKIM and ARC signing keys. New ones can be generated, but each has to be published in DNS again, and mail signed before that fails DKIM. |
|
With |
|
Nothing permanent; certbot obtains new certificates. Keep it anyway if you publish DANE (TLSA) records for the current key. |
|
Only with an OAuth or client-certificate smarthost: the refresh token has to be authorised again. |
The users’ mailboxes (Maildir, or the MDA’s store behind LMTP) are not
Pepsi’s and belong in your ordinary backup of home directories.
Keep the files and the database dump together, and from the same moment: a
dump newer than secrets.d may hold keys wrapped under a key-encryption key
the file backup does not have. Both hold secrets; store them encrypted and
readable by root only.
7.3.2. Taking a backup¶
sudo -u pepsi-owner pg_dump --format=custom pepsi > pepsi-$(date +%F).dump
tar -C / -czpf pepsi-files-$(date +%F).tar.gz \
etc/pepsi var/pepsi var/lib/pepsi-wallets
Run as root, so that tar can read secrets.d. pg_dump takes a
consistent snapshot while the services run; the files change rarely (a new key,
a configuration change), so taking them just before or after the dump is
enough. The database name and its owner are those in [pepsi-postgres].
The dump that pepsi-setup schema --backup-dir takes before an upgrade (see
Backing up before a schema upgrade) covers the database only, and only the pepsi
schema; it does not replace this.
7.3.3. Restoring¶
Restore on a host running the same Pepsi release the backup came from (upgrade afterwards, if you want to), in this order:
Install Pepsi. The Debian package creates the service accounts and database roles; on a source install create them as in Installation.
Stop everything:
systemctl stop pepsi.target.Restore the files, secrets.d first.
tarkeeps owners by name, so the accounts must exist before you unpack:tar -C / -xzpf pepsi-files-2026-10-01.tar.gzRestore the database into an empty one owned by the schema owner:
sudo -u postgres dropdb --if-exists pepsi sudo -u postgres createdb -O pepsi-owner pepsi sudo -u pepsi-owner pg_restore -d pepsi pepsi-2026-10-01.dump
Run
pepsi-setup -c /etc/pepsi/pepsi.conf run. It re-applies the role grants, writes again what lives outside both backups (the systemd drop-ins that hand the TLS certificates to the services) and checks the DNS records against the restored keys. It does not replace a key or secret that already exists.Start the services:
systemctl start pepsi.target.
If pepsi-setup says it refuses to generate a key-encryption key because
identities exist, secrets.d/pepsi-crypto.secret was not restored, or is not
the one that goes with this database. Do not work around it: find the right
file.