26. Benchmark Suite

The tests/NN-*-bench.sh scripts measure how fast the live Pepsi pipeline processes mail. They reuse the same hosts, accounts and helpers as the correctness suite (tests/test-accounts.ini, tests/lib.sh) plus a shared tests/bench-lib.sh, and are driven the same way:

tests/08-perstage-bench.sh tests/test-accounts.ini
tests/09-latency-bench.sh  tests/test-accounts.ini
tests/10-goodput-bench.sh  tests/test-accounts.ini
# or all three in order:
make benchmarks ACCOUNTS=tests/test-accounts.ini

They are not part of make integrationtests (they are slow and deliberately load the MTA). Each one runs preflight_common first, so it SKIPs cleanly if the live hosts/DB/services are unreachable. Run tests/01-deploy.sh first so the pipeline — including the DAVE local-Maildir delivery target — exists.

This chapter describes the method. For how to read the results — per-stage cost, idle latency and saturation goodput — and what to tune, see the Performance chapter.

26.1. How the numbers are obtained

  • Per-stage cost comes from the dispatcher’s own timing: pepsi-dispatch records every stage execution into pepsi.stage_stats(stage, messages, duration_us). Snapshotting that table before/after a message gives exact microseconds-per-message per stage — for a stage reached by an SMTP-sent message and for one injected straight at it alike, since both run as ordinary dispatched workers. The benchmarks lower [pepsi-dispatch] STATS_INTERVAL (and POLL_INTERVAL) so the counters flush promptly, and also set [pepsi] LOG to a quiet level (warn by default, overridable via BENCH_LOG_LEVEL) so per-message logging does not inflate the timings — that knob is read by every component, so it covers the dispatcher and the stage workers it spawns (which take no log flag of their own). Both pepsi-dispatch and pepsi-ingress are restarted so the new level takes effect, and the benchmarks restore the deployed config on exit (a one-shot pepsi.conf.bench.bak backup on the ADMIN host, replayed by a trap … EXIT).

  • Goodput is the delta of the global pepsi.dispatch_stats.messages_processed counter over each 5-second window. It is flushed every STATS_INTERVAL, so at a high rate it lags the truth by up to one interval; the deliv/s column beside it — files appearing in the destination Maildir/new — is the instantaneous cross-check, and the two disagreeing by a step’s worth is that lag, not a lost message.

  • Delivery is cross-checked by counting files in the target Maildir/new.

  • The load generator is several processes, not one. load_burst splits a burst across LOAD_PROCS (default 4) OS processes, each running its share of the threads over its own stride of the message range, and merges their JSON. One Python process is one GIL, and at high rates the interpreter is what stops the number going up — which means a single-process generator stops measuring the system under test and starts measuring itself.

To keep the inbound pay-to-send (anti-spam) stage from gating the load, the benchmarks pre-whitelist the sender (bench_whitelist_sender), so check-whitelist sets state.spam=false and anti-spam short-circuits.

Note

The ramp cannot build a backlog through SMTP, and that is not a flaw in the ramp. pepsi-ingress commits each admission synchronously — a 250 means the message reached the disk — so on a host where admission is the narrower of the two, the queue stays near empty however hard the generator pushes and the ramp measures the front end. To measure the pipeline with a deep queue, insert rows straight into pepsi.workqueue at status = 'pending' in one transaction and time the drain; the Performance chapter shows how, and the difference between the two figures is the interesting part.

26.2. The three benchmarks

  1. ``08-perstage-bench.sh`` — per-stage cost vs message size. Prints one stage × size matrix of µs/msg covering every stage Pepsi has, filled in by two halves that answer two different questions: the scenarios time the stages of the deployed pipeline in situ, and the isolated stages time every stage program the deployed pipeline routes no message through.

    The scenarios. A single message only traverses one path, so it can only time the stages on that path; covering the whole bidirectional, branching pipeline therefore takes several messages — one per path. The benchmark drives a set of SCENARIOS, each a message whose journey reveals a different group of stages, and accumulates their per-stage costs into the matrix:

    scenario

    message

    stages it newly reveals

    inbound-local

    CAROL → DAVE (foreign inbound → local Maildir)

    init (arc), route (if), decrypt, detect-language, block-language, check-whitelist, anti-spam (short-circuit), list (list router), forward (dot-forward), local (maildir)

    inbound-relay

    CAROL → ALICE (inbound → onward relay)

    srs, dkim-sign-relay, smarthost (the non-local recipient is peeled by local onto that tail)

    outbound

    ALICE → CAROL (local-origin submission → direct relay)

    list-out (list router), edit-settings, auto-whitelist, delay-route (if), encrypt, dkim-sign, internet (relay-to-internet)

    bounce

    CAROL → dave-broken (local delivery fails)

    bounce (the maildir helper fails → a bounce SideClone routes to the bounce stage → dkim-sign)

    bounce-language

    CAROL → DAVE, French body (blacklisted language)

    bounce-language (block-language bounces an fr message)

    payment (opt-in)

    CAROL (not whitelisted) → DAVE (pay-to-send gate)

    bounce-payment (shortens PAYMENT_DEADLINE for the run)

    delay (opt-in)

    ALICE → blackhole.invalid, ENVID=PEPSIDELAY (DELAY DSN)

    delaytest (relay-to-smarthost to an unrouteable MTA)

    The default SCENARIOS are the five fast ones; append payment and/or delay to also time the last two deployed stages (they are slower — merchant- and retry-bound). The payment/delay scenarios and the two opt-in stages need a sender that is not local-origin and the broken local target, all of which 01-deploy.sh provisions. Setting SCENARIOS="" runs the isolated half alone.

    The isolated stages. pepsi-ingress always admits a message at init (pepsi_common::stage::INITIAL_STAGE), so a scenario can only ever reach the stages the deployed pipeline routes it through. Every stage program that is not wired into that pipeline is therefore unreachable by any message, however constructed.

    So the run has a second half. It appends benchmark-only [stage-bench-*] sections to the deployed configuration — each with that stage’s own sensible defaults, and every edge pointing at one shared bench-discard sink (pepsi-stage-discard), so nothing an isolated stage advances can escape into the live pipeline or onto the network — and then injects a sized message straight at each of them with the pepsi.workqueue_inject SQL function, which is exactly what a stage that originates mail does (StageContext::enqueue_new). workqueue_inject notifies the workqueue channel, so the dispatcher picks the row up at once, the stage runs as an ordinary worker, and its cost lands in pepsi.stage_stats like any other. Injecting rather than sending also takes pepsi-ingress, SMTP and the admission limits out of the measurement, so an isolated figure is the stage’s own cost and nothing else.

    isolated stage

    program

    what is configured, and what it therefore measures

    bench-discard

    pepsi-stage-discard

    DISPOSITION = success, BOUNCE = no — the shared sink, so its own row is also the pipeline’s bare per-message floor

    bench-aliases

    pepsi-stage-aliases

    a three-entry ALIASES map (a two-target expansion, a transitive hop, an @domain catch-all) and a recipient that expands

    bench-route

    pepsi-stage-route

    one explicit ROUTES carve-out plus MANAGED_STAGE; the recipient is at a managed domain, so the managed branch is taken

    bench-vacation

    pepsi-stage-vacation

    VACATION_RANGES covering today (in the server’s timezone) and SUPPRESS_DAYS = 0, so every message really composes and injects a notice rather than taking the “nobody is away” exit

    bench-autocrypt-learn

    pepsi-stage-autocrypt-learn

    defaults, on a message with no Autocrypt: header — the pass over a message with nothing to learn, which is what all but a sliver of real inbound mail is

    bench-vks-confirm

    pepsi-stage-vks-confirm

    a VKS_HOST the bench sender is not, so the message passes through and no link is followed (no network I/O)

    bench-auto-pay

    pepsi-stage-auto-pay

    the wallet configured as the deployed [stage-pay], on a message that is not a payment demand — again the common inbound case. The paying half needs a live merchant and wallet; 12/13 drive that.

    bench-secure-link

    pepsi-stage-secure-link

    defaults; terminal. The global [pepsi-secure-link] NOTIFY_STAGE is pointed at the sink for the run, so timing the stage puts no mail on the wire.

    bench-lmtp

    pepsi-stage-relay-to-lmtp

    the local Dovecot socket and the throwaway mailbox 11-lmtp-test.sh provisions; skipped unless both are present

    bench-milter

    pepsi-stage-milter

    skipped unless a milter is named in BENCH_MILTER_ADDR (e.g. inet:8890@127.0.0.1)

    An isolated stage whose program is not installed, or whose external dependency is missing, is dropped before the sections are written; and if the dispatcher does not come back up with them, the append is rolled back and the isolated half is abandoned rather than leaving the host without its coordinator. ISO_STAGES="" skips the half entirely.

    Deliberately not isolated: pepsi-stage-encrypt, -decrypt, -dot-forward, -srs, -dkim-sign and -if — the deployed pipeline routes through all of them, so the scenarios already time them in situ, which is the more honest figure.

    Reading the matrix. Body-reading stages (arc verify, decrypt, detect-language, encrypt, dkim-sign, secure-link, relay/maildir) climb with size; metadata-only stages (if-branches, srs, whitelist, aliases, route) stay flat. The relay stages (smarthost, internet, delaytest, bench-lmtp) are network- or MDA-inclusive — they hold the outbound transaction open, so their figure is the next hop’s round-trip, not local CPU; the benchmark prints an explicit note when such a stage ran more times than messages were injected, i.e. when the next hop deferred and the figure averages retries in.

  2. ``09-latency-bench.sh`` — idle end-to-end latency. Runs the whole inject→observe loop on the ADMIN host (one clock, no ssh-per-poll) for N short messages (default 200) and reports min/mean/median/p95/max, plus the sum of mean per-stage processing time as a floor. POLL_MS (default 2) is how often it looks for the delivered message, and therefore the quantisation of every figure it prints, so it has to be small against the latency being measured: a poll interval near the real latency rounds the distribution into a few buckets and lets the poll rather than the pipeline decide the p95.

  3. ``10-goodput-bench.sh`` — saturation goodput, three load paths, two instruments. The three paths are A pure pipeline (local inject → local Maildir), B network + front-end (remote sender → local Maildir) and C outbound relay (authorized ALICE origin → relay-to-internet → CAROL’s MX). Each is measured twice:

    • a fixed-burst drain (run_burst) offers BURST_MSGS messages flat out and times the queue to empty, reporting the rate and the host’s CPU, disk %util and link utilisation over exactly that window;

    • a ramp (run_ramp) doubles the offered load every STEP_SECS until the queue is over SAT_QUEUE and still growing.

    The burst is the figure; the ramp is the shape. A ramp attributes completions to whichever fixed window they land in, and starting a generator costs an ssh round-trip and a Python start-up on the injecting host — a large share of a five-second window when that host is remote, and larger still when it is small, so the ramp reads well below the burst for the same path. Read the ramp for where the queue starts growing and what the host is doing there.

    The burst for each scenario runs before that scenario’s ramp, because a ramp deliberately ends saturated: it leaves a large backlog and a Maildir with as many files in it, and a burst measured on top of that recovery reads far below one on a settled machine.

    A burst aimed at a mailbox on another host uses BURST_MSGS_REMOTE (2 000) rather than BURST_MSGS (20 000), and empties that mailbox afterwards. Both are about the next script rather than this one: tens of thousands of messages relayed to a small peer leave it delivering for a long time, and every helper that finds a message does it by scanning the mailbox, so a peer still holding them makes the next script’s lookups time out even though it delivered every message.

    Four things about how it ramps are worth knowing before reading a number out of it.

    • The ramp is geometric. Each step multiplies the offered load by RAMP_FACTOR (default 2) instead of adding a fixed RATE_UNIT. A linear ramp cannot serve both ends of the range this suite is pointed at: a small VM and a large server plateau orders of magnitude apart, so a unit that finds the small host’s knee stops far short of the big one and reports MAX_STEPS as its ceiling. What it costs is resolution at the knee — the offered rate that provoked the plateau is bracketed only to within a factor of RAMP_FACTOR. The sustained figure is not affected: that is the measured completion rate. RAMP_FACTOR=1 gives a linear ramp of RATE_UNIT steps.

    • Saturation is a growing queue, not a flat rate. The ramp stops when the queue is “over SAT_QUEUE and still growing for STALL_STEPS (default 2) consecutive steps”, which is what saturation is — more admitted than completed. A plateau in the processed rate would be the wrong signal: the per-window rate keeps wobbling either way long after the pipeline is flat out, so a plateau test can run out of steps while the queue is already deep.

    • It reports three utilisations, and only one of them may be extrapolated from. CPU, the %util of the disk holding PGDATA, and the NIC’s link utilisation are sampled over each burst and each ramp step. CPU tracks the load; the link is nowhere near binding but could be on a busier one; disk %util must not be used, because it counts wall time with a non-empty request queue rather than device saturation, and on a parallel NVMe device it need not rise with throughput at all. The summary’s extrapolated column therefore scales by CPU and caps the result at the best rate any all-local path actually achieved (see Extrapolating past a remote bottleneck).

    • It raises the pipeline’s own worker count, not just the admission caps: BENCH_STAGE_PAR (default 16) is written as PARALLELISM into every [stage-*] section for the run. The deployed default of 4 exists so that an out-of-the-box install fits an out-of-the-box max_connections; measuring it and calling the answer the machine’s ceiling is measuring the default. Because each stage worker holds exactly one database connection, the script then prints Σ PARALLELISM plus the dispatcher and server pools against PostgreSQL’s max_connections and warns when it does not fit — a worker that cannot get a connection reports database-overload, the dispatcher requeues and throttles, and the ramp quietly reports a lower number, so an unchecked budget error looks exactly like a capacity figure.

26.3. Rate-limit reporting

The parallel load generator records the SMTP reply of every rejection, so Pepsi’s own front-end throttle ([pepsi-ingress] CONN_RATE_PER_SECOND / CONN_RATE_BURST) shows up as 421s and is reported with the knob to raise.

For the outbound path (scenario C), which touches MTAs Pepsi does not control, the script reads Pepsi’s view of the next hop — state.bounce SMTP codes/text, pepsi.tls_session, the relay journal — and, if it sees 4xx / “rate”/”too many”/”try again”/”throttle” deferrals, reports external rate-limiting with concrete fixes:

  • On the remote MTA (Postfix): raise smtpd_client_connection_rate_limit, smtpd_client_message_rate_limit, the *_connection_count_limit, widen anvil_rate_time_unit, or exempt the Pepsi host’s IP via smtpd_client_event_limit_exceptions / a dedicated transport.

  • On the Pepsi side (to stay under a limit you cannot change): lower [stage-internet] PARALLELISM and pace retries via RETRY_INITIAL / RETRY_MAX_INTERVAL / RETRY_FACTOR.

26.4. Useful knobs

script

knobs

all

BENCH_STATS_INTERVAL, BENCH_POLL_INTERVAL, BENCH_LOG_LEVEL, MSGS_PER_CONN, LOAD_PROCS, BENCH_WHITELIST_NAME, BENCH_CONFIRM_GRACE; and, for the runs that call bench_set_admission, BENCH_CONN_RATE / BENCH_CONN_BURST / BENCH_MAX_CONN / BENCH_DB_POOL / BENCH_RELAY_PAR_MULT / BENCH_QUEUE_LIMIT / BENCH_STAGE_PAR / BENCH_DISPATCH_POOL

08

SIZES, REPEAT, SCENARIOS, PERMSG_TIMEOUT, BOUNCE_TIMEOUT_S, PAY_DEADLINE_S, PAY_TIMEOUT_S; and for the isolated half ISO_STAGES, ISO_SIZES, ISO_REPEAT, ISO_TIMEOUT, ISO_LOCAL_RCPT, ISO_SENDER, ISO_LMTP_SOCKET, ISO_LMTP_USER, BENCH_MILTER_ADDR

09

N, WARMUP, MSG_BYTES, POLL_MS, LATENCY_MAX_MS, PERMSG_TIMEOUT_MS

10

RATE_UNIT, RAMP_FACTOR, PAR_UNIT, PAR_MAX, MAX_STEPS, STEP_SECS, MSG_BYTES, STALL_STEPS, SAT_QUEUE, DRAIN_TIMEOUT, WARMUP_MSGS, WARMUP_DRAIN, RUN_A/RUN_B/RUN_C; and for the burst half BURST_MSGS, BURST_MSGS_REMOTE, BURST_PAR, BURST_TIMEOUT, BURST_SETTLE