8. Troubleshooting

Start with two commands; most problems show up in one of them:

pepsi-status
journalctl -u 'pepsi-*' -p warning --since '1 hour ago'

Operating Pepsi explains the logs and what pepsi-status reports. The problems below are the ones early deployments met most often.

8.1. Common problems

8.1.1. No mail arrives

  • Another MTA holds port 25. Debian installs one by default. systemctl start pepsi.target then fails, and systemctl status pepsi-ingress.socket shows Result: resources. ss -ltnp 'sport = :25' names the program; stop and disable it (see Installation, “Migrating from an existing mail server”).

  • DNS or the firewall. The domain’s MX record has to name this host and port 25 has to be open from the Internet. pepsi-setup -c /etc/pepsi/pepsi.conf run prints the records it expects and checks the published ones.

  • Pepsi refuses the message. The sender’s server gets the reason, and so does the journal of pepsi-ingress. The usual ones: 550 5.1.1 for an address nothing on this host delivers to (see VERIFY_RECIPIENTS in pepsi-ingress(1)); 550 for a definite DMARC failure under DMARC_ENFORCE; 452 4.3.1 when the queue or the disk is full (see MAX_QUEUE_ROWS and MIN_FREE_SPACE in pepsi.conf(5)).

8.1.2. Outgoing mail is not delivered

pepsi-status lists the messages waiting at each stage. Messages that stay paused at a relay stage are being retried; pepsi-queue list ID --state-only shows the last reply of the remote server. Common causes:

  • Outbound port 25 is blocked. Many hosting and residential providers block it. Relay through a smarthost instead (pepsi-setup --wizard).

  • The receiving server rejects or files the mail as spam. Check that the SPF, DKIM and DMARC records pepsi-setup run prints are published, and that the host’s address has a reverse DNS name matching [pepsi-ingress] HOSTNAME.

8.1.3. sendmail or cron mail fails with “temporary failure”

pepsi-sendmail hands mail to pepsi-ingress over /run/pepsi/submission.sock and exits with status 75 (temporary failure) when nobody answers there. Check that pepsi.target is running. If /run/pepsi was deleted by hand, systemd-tmpfiles --create re-creates it with the right permissions; then restart pepsi-ingress.

8.1.4. A program exits with status 78

The database schema is not the one the program was built with, typically in the minutes after an upgrade before pepsi-setup schema has run. The message names the fix; Upgrading has the table of cases. The services retry by themselves and come back once the schema matches.

8.1.5. “permission denied” from the database

The database roles’ grants are set by pepsi-setup. After restoring a database, or after changing roles by hand, run pepsi-setup -c /etc/pepsi/pepsi.conf run again.

8.1.6. Messages in failed or timeout

Most errors do not end a message at all, because they are this host’s fault rather than the message’s:

  • A lost database connection: the message is paused and retried every 30 seconds until the database is back, and each retry is logged under pepsi-dispatch (“lost its database connection”).

  • Any other error a stage reports — a key, template, map or trust store that cannot be read, a helper that cannot be started, a database statement that is refused, a service that does not answer: the message stays paused at its stage with the error in state.last_error, and is retried after one minute, then at doubling intervals up to an hour, for the stage’s MAX_LIFETIME (120 hours unless configured). The stage logs each retry (“failed at stage ‘…’ (attempt N); retrying in …s”). Fix the cause and the queue drains by itself; pepsi-status shows the waiting messages meanwhile. When MAX_LIFETIME runs out the message goes to the stage’s BOUNCE_STAGE, whose DSN tells the sender that a local error persisted.

  • A stage whose configuration no longer parses does not start at all: the dispatcher logs “the worker refused to start”, keeps the messages queued and tries again every few seconds. The stage’s own log line says what is wrong.

  • A worker that crashes or hangs past MAX_RUNTIME pauses the message it was working on and retries it a minute later; only the third time is the message given up on.

  • A milter that keeps deferring a message, or cannot be reached, holds it for MAX_LIFETIME as well, and then bounces it (or leaves it failed) — it never drops it through the filter’s REJECT_STAGE. A milter that defers only some recipients holds just their copy, as a separate paused message at the milter stage, while the others are delivered.

A message ends in failed only when a stage found a defect in the message itself, when a stage with no BOUNCE_STAGE ran out of retries, or when it crashed its worker three times; and in timeout when it hung it three times. state.failure_class and state.last_error say which and why.

Delivery errors from the next server (a refusing or unreachable MX, a DNS lookup timing out) are retried by the relay stages themselves in the same way. A failed message is still often worth retrying once the cause is fixed.

pepsi.target runs pepsi-failure-bouncer --once every ten minutes (pepsi-failure-bouncer.timer). It moves every failed or timeout message that has been so for [pepsi-failure-bouncer] MIN_AGE (an hour unless configured) to the stage named by [pepsi-failure-bouncer] BOUNCE_STAGE (bounce in the shipped and the wizard’s configuration), which sends the sender a delivery-status notification as their NOTIFY asks, and logs each one with the reason. A sender whose domain the message did not authenticate gets no notification under BOUNCE_UNAUTHENTICATED = drop, and the message is deleted, with a warning in the bounce stage’s log. A copy of a mailing-list post is deleted rather than bounced. So a failed message stays in the queue for about MIN_AGE; to retry it rather than bounce it, act within that time:

pepsi-queue list --status failed          # which messages, at which stage
pepsi-queue list 1234 --state-only        # why: state.last_error, state.bounce, …
pepsi-queue set-stage 1234 relay          # retry it from a stage of your choice
pepsi-failure-bouncer --once              # or bounce all of them now
pepsi-queue delete 1234                   # or discard it

The options of list go after the word list. systemctl list-timers pepsi-failure-bouncer.timer shows when the next sweep runs, and journalctl -u pepsi-failure-bouncer how many messages each one moved.

8.1.7. Keeping failed messages for diagnosis

On a development or test host a failed message is usually the thing you want to look at, and a bounce an hour later destroys it. Raise MIN_AGE, or turn the sweep off there:

systemctl mask --now pepsi-failure-bouncer.timer

mask rather than disable: pepsi.target wants the timer, so a merely disabled or stopped timer starts again with the target. Failed messages then stay in the queue until you retry, bounce or delete them as above. To turn the sweep back on:

systemctl unmask pepsi-failure-bouncer.timer
systemctl start pepsi-failure-bouncer.timer

The next sweep then bounces everything that failed in the meantime; delete what you do not want bounced first. Do not leave the timer masked on a server that handles real mail: nobody is told about a message that failed, neither the sender nor you, unless you watch problem_total (see Monitoring).

pepsi-failure-bouncer can instead run as a service that bounces each failed message as it happens. No unit is shipped for that mode; mask the timer before starting it from your own unit.

8.2. Frequently asked questions

Does Pepsi keep a log of who mailed whom?

Not by default. A delivered message is deleted from the queue and only the journal mentions it. [pepsi] MAIL_LOG turns on a per-message record; read the warning in pepsi.conf(5) first.

Where is the queue?

In PostgreSQL, table pepsi.workqueue. pepsi-queue and pepsi-status are the supported ways to look at it; plain SELECT works too.

How do I see the pipeline this host runs?

pepsi-config dump prints the effective configuration (--origin also shows which settings the database overrides), and pepsi-setup run prints, for each domain, how recipients are checked. The Wizard explains the pipelines the wizard builds.

Can I go back to the previous release?

Not in place: the schema is migrated forward only. Restore the backup taken before the upgrade, then install the older packages; see Upgrading.

Can Pepsi run next to Postfix or Exim?

Not both on port 25. Pepsi can import the other server’s configuration; see Installation.

Why does the console show everything read-only?

The pepsi-httpd-admin package, which carries out the console’s changes, is not installed or its socket is not running; see Debian packages.

8.3. Uninstalling

Debian packages. apt remove pepsi stops the services and removes the programs; configuration, keys, the database and the service accounts stay, so a later install carries on where it left off. apt purge pepsi removes the rest of what the package installed but, on purpose, still keeps the configuration, the database, the keys, the wallets and the accounts, and prints the commands that remove them. They are, as root:

runuser -u postgres -- dropdb pepsi
rm -rf /var/pepsi /var/lib/pepsi-wallets /etc/pepsi
for u in pepsi pepsi-ingress pepsi-httpd pepsi-owner pepsi-config \
         pepsi-helper-token-refresh pepsi-keydisc pepsi-crypto \
         pepsi-whitelist pepsi-wallets; do deluser "$u"; done
for g in pepsi-maildir pepsi-forward pepsi-token pepsi-admin \
         pepsi-telemetry; do delgroup "$g"; done

The database roles of the same names can then be dropped with dropuser. Remove pepsi-telemetry and pepsi-httpd-admin the same way if they are installed.

Source installation. Stop and disable pepsi.target, then run make uninstall with the same PREFIX/DESTDIR you installed with. It removes the programs, the shipped data and the systemd units, not the configuration, keys or database; remove those as above.

Warning

Everything removed this way is gone for good, including the key-encryption key without which the stored private keys cannot be opened. Take a backup first (Backup and restore) if there is any chance you will want the keys, the mailing lists or their archive again.