8. Troubleshooting¶
Start with two commands; most problems show up in one of them:
pepsi-status
journalctl -u 'pepsi-*' -p warning --since '1 hour ago'
Operating Pepsi explains the logs and what pepsi-status reports. The
problems below are the ones early deployments met most often.
8.1. Common problems¶
8.1.1. No mail arrives¶
Another MTA holds port 25. Debian installs one by default.
systemctl start pepsi.targetthen fails, andsystemctl status pepsi-ingress.socketshowsResult: resources.ss -ltnp 'sport = :25'names the program; stop and disable it (see Installation, “Migrating from an existing mail server”).DNS or the firewall. The domain’s
MXrecord has to name this host and port 25 has to be open from the Internet.pepsi-setup -c /etc/pepsi/pepsi.conf runprints the records it expects and checks the published ones.Pepsi refuses the message. The sender’s server gets the reason, and so does the journal of
pepsi-ingress. The usual ones:550 5.1.1for an address nothing on this host delivers to (seeVERIFY_RECIPIENTSin pepsi-ingress(1));550for a definite DMARC failure underDMARC_ENFORCE;452 4.3.1when the queue or the disk is full (seeMAX_QUEUE_ROWSandMIN_FREE_SPACEin pepsi.conf(5)).
8.1.2. Outgoing mail is not delivered¶
pepsi-status lists the messages waiting at each stage. Messages that stay
paused at a relay stage are being retried; pepsi-queue list ID
--state-only shows the last reply of the remote server. Common causes:
Outbound port 25 is blocked. Many hosting and residential providers block it. Relay through a smarthost instead (
pepsi-setup --wizard).The receiving server rejects or files the mail as spam. Check that the SPF, DKIM and DMARC records
pepsi-setup runprints are published, and that the host’s address has a reverse DNS name matching[pepsi-ingress] HOSTNAME.
8.1.3. sendmail or cron mail fails with “temporary failure”¶
pepsi-sendmail hands mail to pepsi-ingress over
/run/pepsi/submission.sock and exits with status 75 (temporary failure) when
nobody answers there. Check that pepsi.target is running. If
/run/pepsi was deleted by hand, systemd-tmpfiles --create re-creates it
with the right permissions; then restart pepsi-ingress.
8.1.4. A program exits with status 78¶
The database schema is not the one the program was built with, typically in the
minutes after an upgrade before pepsi-setup schema has run. The message
names the fix; Upgrading has the table of cases. The services retry by
themselves and come back once the schema matches.
8.1.5. “permission denied” from the database¶
The database roles’ grants are set by pepsi-setup. After restoring a
database, or after changing roles by hand, run pepsi-setup -c
/etc/pepsi/pepsi.conf run again.
8.1.6. Messages in failed or timeout¶
Most errors do not end a message at all, because they are this host’s fault rather than the message’s:
A lost database connection: the message is paused and retried every 30 seconds until the database is back, and each retry is logged under
pepsi-dispatch(“lost its database connection”).Any other error a stage reports — a key, template, map or trust store that cannot be read, a helper that cannot be started, a database statement that is refused, a service that does not answer: the message stays
pausedat its stage with the error instate.last_error, and is retried after one minute, then at doubling intervals up to an hour, for the stage’sMAX_LIFETIME(120 hours unless configured). The stage logs each retry (“failed at stage ‘…’ (attempt N); retrying in …s”). Fix the cause and the queue drains by itself;pepsi-statusshows the waiting messages meanwhile. WhenMAX_LIFETIMEruns out the message goes to the stage’sBOUNCE_STAGE, whose DSN tells the sender that a local error persisted.A stage whose configuration no longer parses does not start at all: the dispatcher logs “the worker refused to start”, keeps the messages queued and tries again every few seconds. The stage’s own log line says what is wrong.
A worker that crashes or hangs past
MAX_RUNTIMEpauses the message it was working on and retries it a minute later; only the third time is the message given up on.A milter that keeps deferring a message, or cannot be reached, holds it for
MAX_LIFETIMEas well, and then bounces it (or leaves itfailed) — it never drops it through the filter’sREJECT_STAGE. A milter that defers only some recipients holds just their copy, as a separatepausedmessage at the milter stage, while the others are delivered.
A message ends in failed only when a stage found a defect in the message
itself, when a stage with no BOUNCE_STAGE ran out of retries, or when it
crashed its worker three times; and in timeout when it hung it three times.
state.failure_class and state.last_error say which and why.
Delivery errors from the next server (a refusing or unreachable MX, a DNS lookup timing out) are retried by the relay stages themselves in the same way. A failed message is still often worth retrying once the cause is fixed.
pepsi.target runs pepsi-failure-bouncer --once every ten minutes
(pepsi-failure-bouncer.timer). It moves every failed or timeout
message that has been so for [pepsi-failure-bouncer] MIN_AGE (an hour
unless configured) to the stage named by [pepsi-failure-bouncer]
BOUNCE_STAGE (bounce in the shipped and the wizard’s configuration), which
sends the sender a delivery-status notification as their NOTIFY asks, and
logs each one with the reason. A sender whose domain the message did not
authenticate gets no notification under BOUNCE_UNAUTHENTICATED = drop, and
the message is deleted, with a warning in the bounce stage’s log. A copy of a
mailing-list post is deleted rather than bounced. So a failed message stays in
the queue for about MIN_AGE; to retry it rather than bounce it, act within
that time:
pepsi-queue list --status failed # which messages, at which stage
pepsi-queue list 1234 --state-only # why: state.last_error, state.bounce, …
pepsi-queue set-stage 1234 relay # retry it from a stage of your choice
pepsi-failure-bouncer --once # or bounce all of them now
pepsi-queue delete 1234 # or discard it
The options of list go after the word list.
systemctl list-timers pepsi-failure-bouncer.timer shows when the next sweep
runs, and journalctl -u pepsi-failure-bouncer how many messages each one
moved.
8.1.7. Keeping failed messages for diagnosis¶
On a development or test host a failed message is usually the thing you want to
look at, and a bounce an hour later destroys it. Raise MIN_AGE, or turn the
sweep off there:
systemctl mask --now pepsi-failure-bouncer.timer
mask rather than disable: pepsi.target wants the timer, so a merely
disabled or stopped timer starts again with the target. Failed messages then
stay in the queue until you retry, bounce or delete them as above. To turn the
sweep back on:
systemctl unmask pepsi-failure-bouncer.timer
systemctl start pepsi-failure-bouncer.timer
The next sweep then bounces everything that failed in the meantime; delete what
you do not want bounced first. Do not leave the timer masked on a server that
handles real mail: nobody is told about a message that failed, neither the
sender nor you, unless you watch problem_total (see
Monitoring).
pepsi-failure-bouncer can instead run as a service that bounces each failed message as it happens. No unit is shipped for that mode; mask the timer before starting it from your own unit.
8.2. Frequently asked questions¶
- Does Pepsi keep a log of who mailed whom?
Not by default. A delivered message is deleted from the queue and only the journal mentions it.
[pepsi] MAIL_LOGturns on a per-message record; read the warning in pepsi.conf(5) first.- Where is the queue?
In PostgreSQL, table
pepsi.workqueue.pepsi-queueandpepsi-statusare the supported ways to look at it; plainSELECTworks too.- How do I see the pipeline this host runs?
pepsi-config dumpprints the effective configuration (--originalso shows which settings the database overrides), andpepsi-setup runprints, for each domain, how recipients are checked. The Wizard explains the pipelines the wizard builds.- Can I go back to the previous release?
Not in place: the schema is migrated forward only. Restore the backup taken before the upgrade, then install the older packages; see Upgrading.
- Can Pepsi run next to Postfix or Exim?
Not both on port 25. Pepsi can import the other server’s configuration; see Installation.
- Why does the console show everything read-only?
The
pepsi-httpd-adminpackage, which carries out the console’s changes, is not installed or its socket is not running; see Debian packages.
8.3. Uninstalling¶
Debian packages. apt remove pepsi stops the services and removes the
programs; configuration, keys, the database and the service accounts stay, so a
later install carries on where it left off. apt purge pepsi removes the
rest of what the package installed but, on purpose, still keeps the
configuration, the database, the keys, the wallets and the accounts, and
prints the commands that remove them.
They are, as root:
runuser -u postgres -- dropdb pepsi
rm -rf /var/pepsi /var/lib/pepsi-wallets /etc/pepsi
for u in pepsi pepsi-ingress pepsi-httpd pepsi-owner pepsi-config \
pepsi-helper-token-refresh pepsi-keydisc pepsi-crypto \
pepsi-whitelist pepsi-wallets; do deluser "$u"; done
for g in pepsi-maildir pepsi-forward pepsi-token pepsi-admin \
pepsi-telemetry; do delgroup "$g"; done
The database roles of the same names can then be dropped with dropuser.
Remove pepsi-telemetry and pepsi-httpd-admin the same way if they are
installed.
Source installation. Stop and disable pepsi.target, then run make
uninstall with the same PREFIX/DESTDIR you installed with. It removes
the programs, the shipped data and the systemd units, not the configuration,
keys or database; remove those as above.
Warning
Everything removed this way is gone for good, including the key-encryption key without which the stored private keys cannot be opened. Take a backup first (Backup and restore) if there is any chance you will want the keys, the mailing lists or their archive again.