85.1.51. pepsi-failure-bouncer

Funnel failed/timeout messages to the bounce stage

Manual section:

1

85.1.51.1.1. Name

pepsi-failure-bouncer - move stuck (failed/timeout) messages to the bounce stage.

85.1.51.1.2. Synopsis

pepsi-failure-bouncer [GLOBAL-OPTIONS] [–failed-only | –timeout-only]

pepsi-failure-bouncer [GLOBAL-OPTIONS] –once [–failed-only | –timeout-only]

85.1.51.1.3. Description

A message moving through the post-ingress stage pipeline (see pepsi-dispatch(1)) lands in one of two terminal states when its current stage gives up:

failed

A stage called fail() or returned an error it marked permanent, the worker gave up retrying an error at a stage with no BOUNCE_STAGE, or the message crashed its worker three times.

timeout

The message held its worker past [pepsi-dispatch] MAX_RUNTIME three times.

Most errors never get here: the worker retries them until the stage’s MAX_LIFETIME and then hands the message to the stage’s BOUNCE_STAGE itself (see pepsi-dispatch(1)).

The dispatcher never re-queues a failed or timeout row, so without this program such messages accumulate in pepsi.workqueue and the original sender is never told their mail was not delivered. That is why pepsi.target runs it (see Systemd below).

pepsi-failure-bouncer moves each stuck message to the configured [pepsi-failure-bouncer] BOUNCE_STAGE and resets its status to pending. The dispatcher then runs the bounce stage (pepsi-stage-bounce(1)), which — honouring the sender’s RFC 3461 NOTIFY — turns the message into a delivery-status notification and relays it back to the sender.

It waits first. A message is moved only once it has been failed for [pepsi-failure-bouncer] MIN_AGE (default one hour, measured from state.failed_at, which the database stamps when the row becomes failed/timeout): the operator’s window to notice the failure and put the message back with pepsi-queue(1) before its sender is told. Each message it moves is logged as a warning with its sender, the stage it failed at and the recorded error, because the bounce stage replaces that record with the DSN. The stage is also kept as state.failed_stage, which pepsi-stage-bounce(1) names in the DSN.

A mailing-list copy (a message carrying state.list) is deleted instead, with the same log line. Its envelope sender is the list’s own bounce address, so a DSN would reach the list’s bounce processing and count a fault of this host against the member.

A message that is already at the bounce stage is left untouched, so a bounce that itself fails at that stage does not loop. A bounce that fails downstream of the bounce stage (a null-sender DSN that could not be relayed) is still moved back to the bounce stage, which simply drops it — a bounce is never re-bounced.

The first rule is only the one-step guard; the second ends every longer chain, because a real bounce stage rewrites the message to a null sender and a null-sender message is dropped rather than bounced again. So BOUNCE_STAGE must lead to a bounce stage. Point it at a stage that merely advances and there is no fixed point at all — the message advances, fails again further down, and is reset again, in a loop the failure notification drives at database speed. pepsi-setup(1) fails outright when BOUNCE_STAGE names no configured stage at all, and warns when it names one that does not run pepsi-stage-bounce(1) directly; the second is only a warning because the option may deliberately name a routing stage that reaches the bounce stage on a branch.

Having moved rows to pending, the bouncer notifies the workqueue channel itself, so the dispatcher processes the bounce immediately rather than at its next POLL_INTERVAL sweep. It has to: it runs outside the dispatcher’s loop, and pepsi.workqueue carries no notifying trigger (see pepsi-dispatch(1)).

It connects to the shared database through the [pepsi-postgres] section and is configured by [pepsi-failure-bouncer] (see pepsi.conf(5)). When started as root — a root shell, a cron job, a unit without a User= — it continues as the unprivileged pepsi service account, the role that owns the queue, before connecting: root has no PostgreSQL role of its own, so without that switch the connection would be refused. The configuration is read first, and if the pepsi account does not exist the identity is left untouched and a warning is logged. It is not a stage.

85.1.51.1.4. Modes

With no --once flag, pepsi-failure-bouncer runs as a long-lived service. It LISTENs on the PostgreSQL workqueue_failed channel — fired by the workqueue_failed_notify trigger whenever any writer makes a row failed or timeout — and bounces each message once it is MIN_AGE old, waking for the oldest waiting failure when no notification comes first. On start-up, and again on every reconnect, it first sweeps all messages that are already stuck, so nothing committed before the listener attached (or during a dropped connection) is missed. It runs until it receives SIGINT or SIGTERM.

With –once, it performs that sweep a single time over all currently-stuck messages and exits — the form for a manual operator run and for the shipped timer.

85.1.51.1.5. Systemd

pepsi-failure-bouncer.timer runs pepsi-failure-bouncer.service, a –once sweep as the pepsi account, ten minutes after boot and every ten minutes after that. pepsi.target wants the timer, so enabling the target enables the sweep. A failed message therefore stays in the queue for between MIN_AGE and MIN_AGE plus ten minutes before it is bounced (or, for a sender the message did not authenticate, dropped with a warning; see BOUNCE_UNAUTHENTICATED in pepsi.conf(5)). To retry a message instead, use pepsi-queue(1) set-stage within that time.

Developers and testers who want failed messages kept for diagnosis turn the sweep off with:

systemctl mask --now pepsi-failure-bouncer.timer

Masking is needed because pepsi.target would start a merely disabled or stopped timer again. systemctl unmask pepsi-failure-bouncer.timer followed by systemctl start pepsi-failure-bouncer.timer turns it back on; the next sweep then bounces everything that failed meanwhile. Do not leave it masked on a host that handles real mail.

No unit runs the long-lived service mode. A site that prefers it (a failure is then bounced within moments rather than minutes) masks the timer and starts the service from a unit of its own.

85.1.51.1.6. Options

–once

Sweep the messages that are stuck now, and have been for MIN_AGE, and exit, instead of running as a service that listens for new failures.

–failed-only

Act only on failed messages, ignoring timeout.

–timeout-only

Act only on timeout messages, ignoring failed. Mutually exclusive with –failed-only.

85.1.51.1.7. Global Options

These options may appear before or after the other flags.

-c FILE, –config FILE

Read the configuration from FILE instead of searching the default locations.

-L LOGLEVEL, –log LOGLEVEL

Set the logging verbosity. LOGLEVEL is one of error, warn, info, debug or trace (default: info).

-v, –verbose

Show log messages from all sources, including third-party libraries.

-h, –help

Print a usage summary and exit.

-V, –version

Print the version and exit.

85.1.51.1.8. Exit Status

0

Successful completion (a --once sweep finished, or the service shut down cleanly on a signal).

1

An error occurred: a malformed configuration, the missing required BOUNCE_STAGE option, or a failed database connection. In service mode a lost database connection is not fatal — it is logged and retried with back-off.

85.1.51.1.9. Examples

Run as a service, from a unit of your own:

pepsi-failure-bouncer -c /etc/pepsi/pepsi.conf

Bounce everything currently stuck, once, by hand:

pepsi-failure-bouncer -c /etc/pepsi/pepsi.conf --once

Bounce only the messages that timed out:

pepsi-failure-bouncer -c /etc/pepsi/pepsi.conf --once --timeout-only

85.1.51.1.10. See Also

pepsi-dispatch(1), pepsi-stage-bounce(1), pepsi-queue(1), pepsi-status(1), pepsi.conf(5), systemd.timer(5)

85.1.51.1.11. Bugs

Report bugs to the Pepsi issue tracker.