85.1.51. pepsi-failure-bouncer¶
Funnel failed/timeout messages to the bounce stage
- Manual section:
1
85.1.51.1.1. Name¶
pepsi-failure-bouncer - move stuck (failed/timeout) messages to the bounce stage.
85.1.51.1.2. Synopsis¶
pepsi-failure-bouncer [GLOBAL-OPTIONS] [–failed-only | –timeout-only]
pepsi-failure-bouncer [GLOBAL-OPTIONS] –once [–failed-only | –timeout-only]
85.1.51.1.3. Description¶
A message moving through the post-ingress stage pipeline (see pepsi-dispatch(1)) lands in one of two terminal states when its current stage gives up:
failedA stage called
fail()or returned an error it marked permanent, the worker gave up retrying an error at a stage with noBOUNCE_STAGE, or the message crashed its worker three times.timeoutThe message held its worker past
[pepsi-dispatch] MAX_RUNTIMEthree times.
Most errors never get here: the worker retries them until the stage’s
MAX_LIFETIME and then hands the message to the stage’s BOUNCE_STAGE
itself (see pepsi-dispatch(1)).
The dispatcher never re-queues a failed or timeout row, so without
this program such messages accumulate in pepsi.workqueue and the original
sender is never told their mail was not delivered. That is why
pepsi.target runs it (see Systemd below).
pepsi-failure-bouncer moves each stuck message to the configured
[pepsi-failure-bouncer] BOUNCE_STAGE and resets its status to
pending. The dispatcher then runs the bounce stage
(pepsi-stage-bounce(1)), which — honouring the sender’s RFC 3461 NOTIFY
— turns the message into a delivery-status notification and relays it back to the
sender.
It waits first. A message is moved only once it has been failed for
[pepsi-failure-bouncer] MIN_AGE (default one hour, measured from
state.failed_at, which the database stamps when the row becomes
failed/timeout): the operator’s window to notice the failure and put the
message back with pepsi-queue(1) before its sender is told. Each message it
moves is logged as a warning with its sender, the stage it failed at and the
recorded error, because the bounce stage replaces that record with the DSN. The
stage is also kept as state.failed_stage, which pepsi-stage-bounce(1)
names in the DSN.
A mailing-list copy (a message carrying state.list) is deleted instead,
with the same log line. Its envelope sender is the list’s own bounce address, so a
DSN would reach the list’s bounce processing and count a fault of this host
against the member.
A message that is already at the bounce stage is left untouched, so a bounce that itself fails at that stage does not loop. A bounce that fails downstream of the bounce stage (a null-sender DSN that could not be relayed) is still moved back to the bounce stage, which simply drops it — a bounce is never re-bounced.
The first rule is only the one-step guard; the second ends every longer chain,
because a real bounce stage rewrites the message to a null sender and a
null-sender message is dropped rather than bounced again. So BOUNCE_STAGE
must lead to a bounce stage. Point it at a stage that merely advances and
there is no fixed point at all — the message advances, fails again further down,
and is reset again, in a loop the failure notification drives at database speed.
pepsi-setup(1) fails outright when BOUNCE_STAGE names no configured
stage at all, and warns when it names one that does not run
pepsi-stage-bounce(1) directly; the second is only a warning because the
option may deliberately name a routing stage that reaches the bounce stage on a
branch.
Having moved rows to pending, the bouncer notifies the workqueue channel
itself, so the dispatcher processes the bounce immediately rather than at its next
POLL_INTERVAL sweep. It has to: it runs outside the dispatcher’s loop, and
pepsi.workqueue carries no notifying trigger (see pepsi-dispatch(1)).
It connects to the shared database through the [pepsi-postgres] section and
is configured by [pepsi-failure-bouncer] (see pepsi.conf(5)). When
started as root — a root shell, a cron job, a unit without a User= — it
continues as the unprivileged pepsi service account, the role that owns the
queue, before connecting: root has no PostgreSQL role of its own, so without
that switch the connection would be refused. The configuration is read first, and
if the pepsi account does not exist the identity is left untouched and a
warning is logged. It is not a stage.
85.1.51.1.4. Modes¶
With no --once flag, pepsi-failure-bouncer runs as a long-lived
service. It LISTENs on the PostgreSQL workqueue_failed channel — fired
by the workqueue_failed_notify trigger whenever any writer makes a row
failed or timeout — and bounces each message once it is MIN_AGE old, waking for the oldest
waiting failure when no notification comes first. On start-up,
and again on every reconnect, it first sweeps all messages that are already stuck,
so nothing committed before the listener attached (or during a dropped
connection) is missed. It runs until it receives SIGINT or SIGTERM.
With –once, it performs that sweep a single time over all currently-stuck messages and exits — the form for a manual operator run and for the shipped timer.
85.1.51.1.5. Systemd¶
pepsi-failure-bouncer.timer runs pepsi-failure-bouncer.service, a
–once sweep as the pepsi account, ten minutes after boot and every ten
minutes after that. pepsi.target wants the timer, so enabling the target
enables the sweep. A failed message therefore stays in the queue for between
MIN_AGE and MIN_AGE plus ten minutes before it is bounced (or, for a
sender the message did not authenticate, dropped with a warning; see
BOUNCE_UNAUTHENTICATED in pepsi.conf(5)). To retry a message instead,
use pepsi-queue(1) set-stage within that time.
Developers and testers who want failed messages kept for diagnosis turn the sweep off with:
systemctl mask --now pepsi-failure-bouncer.timer
Masking is needed because pepsi.target would start a merely disabled or
stopped timer again. systemctl unmask pepsi-failure-bouncer.timer followed
by systemctl start pepsi-failure-bouncer.timer turns it back on; the next
sweep then bounces everything that failed meanwhile. Do not leave it masked on
a host that handles real mail.
No unit runs the long-lived service mode. A site that prefers it (a failure is then bounced within moments rather than minutes) masks the timer and starts the service from a unit of its own.
85.1.51.1.6. Options¶
- –once
Sweep the messages that are stuck now, and have been for MIN_AGE, and exit, instead of running as a service that listens for new failures.
- –failed-only
Act only on
failedmessages, ignoringtimeout.- –timeout-only
Act only on
timeoutmessages, ignoringfailed. Mutually exclusive with –failed-only.
85.1.51.1.7. Global Options¶
These options may appear before or after the other flags.
- -c FILE, –config FILE
Read the configuration from FILE instead of searching the default locations.
- -L LOGLEVEL, –log LOGLEVEL
Set the logging verbosity. LOGLEVEL is one of
error,warn,info,debugortrace(default:info).- -v, –verbose
Show log messages from all sources, including third-party libraries.
- -h, –help
Print a usage summary and exit.
- -V, –version
Print the version and exit.
85.1.51.1.8. Exit Status¶
- 0
Successful completion (a
--oncesweep finished, or the service shut down cleanly on a signal).- 1
An error occurred: a malformed configuration, the missing required
BOUNCE_STAGEoption, or a failed database connection. In service mode a lost database connection is not fatal — it is logged and retried with back-off.
85.1.51.1.9. Examples¶
Run as a service, from a unit of your own:
pepsi-failure-bouncer -c /etc/pepsi/pepsi.conf
Bounce everything currently stuck, once, by hand:
pepsi-failure-bouncer -c /etc/pepsi/pepsi.conf --once
Bounce only the messages that timed out:
pepsi-failure-bouncer -c /etc/pepsi/pepsi.conf --once --timeout-only
85.1.51.1.10. See Also¶
pepsi-dispatch(1), pepsi-stage-bounce(1), pepsi-queue(1), pepsi-status(1), pepsi.conf(5), systemd.timer(5)
85.1.51.1.11. Bugs¶
Report bugs to the Pepsi issue tracker.