26. Benchmark Suite¶
The tests/NN-*-bench.sh scripts measure how fast the live Pepsi pipeline
processes mail. They reuse the same hosts, accounts and helpers as the
correctness suite (tests/test-accounts.ini,
tests/lib.sh) plus a shared tests/bench-lib.sh, and are driven the same
way:
tests/08-perstage-bench.sh tests/test-accounts.ini
tests/09-latency-bench.sh tests/test-accounts.ini
tests/10-goodput-bench.sh tests/test-accounts.ini
# or all three in order:
make benchmarks ACCOUNTS=tests/test-accounts.ini
They are not part of make integrationtests (they are slow and
deliberately load the MTA). Each one runs preflight_common first, so it SKIPs
cleanly if the live hosts/DB/services are unreachable. Run tests/01-deploy.sh
first so the pipeline — including the DAVE local-Maildir delivery target —
exists.
This chapter describes the method. For how to read the results — per-stage cost, idle latency and saturation goodput — and what to tune, see the Performance chapter.
26.1. How the numbers are obtained¶
Per-stage cost comes from the dispatcher’s own timing:
pepsi-dispatchrecords every stage execution intopepsi.stage_stats(stage, messages, duration_us). Snapshotting that table before/after a message gives exact microseconds-per-message per stage — for a stage reached by an SMTP-sent message and for one injected straight at it alike, since both run as ordinary dispatched workers. The benchmarks lower[pepsi-dispatch] STATS_INTERVAL(andPOLL_INTERVAL) so the counters flush promptly, and also set[pepsi] LOGto a quiet level (warnby default, overridable viaBENCH_LOG_LEVEL) so per-message logging does not inflate the timings — that knob is read by every component, so it covers the dispatcher and the stage workers it spawns (which take no log flag of their own). Bothpepsi-dispatchandpepsi-ingressare restarted so the new level takes effect, and the benchmarks restore the deployed config on exit (a one-shotpepsi.conf.bench.bakbackup on the ADMIN host, replayed by atrap … EXIT).Goodput is the delta of the global
pepsi.dispatch_stats.messages_processedcounter over each 5-second window. It is flushed everySTATS_INTERVAL, so at a high rate it lags the truth by up to one interval; thedeliv/scolumn beside it — files appearing in the destinationMaildir/new— is the instantaneous cross-check, and the two disagreeing by a step’s worth is that lag, not a lost message.Delivery is cross-checked by counting files in the target
Maildir/new.The load generator is several processes, not one.
load_burstsplits a burst acrossLOAD_PROCS(default 4) OS processes, each running its share of the threads over its own stride of the message range, and merges their JSON. One Python process is one GIL, and at high rates the interpreter is what stops the number going up — which means a single-process generator stops measuring the system under test and starts measuring itself.
To keep the inbound pay-to-send (anti-spam) stage from gating the load, the
benchmarks pre-whitelist the sender (bench_whitelist_sender), so
check-whitelist sets state.spam=false and anti-spam short-circuits.
Note
The ramp cannot build a backlog through SMTP, and that is not a flaw in the
ramp. pepsi-ingress commits each admission synchronously — a 250
means the message reached the disk — so on a host where admission is the
narrower of the two, the queue stays near empty however hard the generator
pushes and the ramp measures the front end. To measure the pipeline with a
deep queue, insert rows straight into pepsi.workqueue at status =
'pending' in one transaction and time the drain; the Performance chapter shows how, and the difference between the two figures
is the interesting part.
26.2. The three benchmarks¶
``08-perstage-bench.sh`` — per-stage cost vs message size. Prints one
stage × sizematrix of µs/msg covering every stage Pepsi has, filled in by two halves that answer two different questions: the scenarios time the stages of the deployed pipeline in situ, and the isolated stages time every stage program the deployed pipeline routes no message through.The scenarios. A single message only traverses one path, so it can only time the stages on that path; covering the whole bidirectional, branching pipeline therefore takes several messages — one per path. The benchmark drives a set of
SCENARIOS, each a message whose journey reveals a different group of stages, and accumulates their per-stage costs into the matrix:scenario
message
stages it newly reveals
inbound-localCAROL→DAVE(foreign inbound → local Maildir)init(arc),route(if),decrypt,detect-language,block-language,check-whitelist,anti-spam(short-circuit),list(list router),forward(dot-forward),local(maildir)inbound-relayCAROL→ALICE(inbound → onward relay)srs,dkim-sign-relay,smarthost(the non-local recipient is peeled bylocalonto that tail)outboundALICE→CAROL(local-origin submission → direct relay)list-out(list router),edit-settings,auto-whitelist,delay-route(if),encrypt,dkim-sign,internet(relay-to-internet)bounceCAROL→dave-broken(local delivery fails)bounce(the maildir helper fails → a bounceSideCloneroutes to the bounce stage →dkim-sign)bounce-languageCAROL→DAVE, French body (blacklisted language)bounce-language(block-languagebounces anfrmessage)payment(opt-in)CAROL(not whitelisted) →DAVE(pay-to-send gate)bounce-payment(shortensPAYMENT_DEADLINEfor the run)delay(opt-in)ALICE→blackhole.invalid,ENVID=PEPSIDELAY(DELAY DSN)delaytest(relay-to-smarthost to an unrouteable MTA)The default
SCENARIOSare the five fast ones; appendpaymentand/ordelayto also time the last two deployed stages (they are slower — merchant- and retry-bound). Thepayment/delayscenarios and the two opt-in stages need a sender that is not local-origin and the broken local target, all of which01-deploy.shprovisions. SettingSCENARIOS=""runs the isolated half alone.The isolated stages.
pepsi-ingressalways admits a message atinit(pepsi_common::stage::INITIAL_STAGE), so a scenario can only ever reach the stages the deployed pipeline routes it through. Every stage program that is not wired into that pipeline is therefore unreachable by any message, however constructed.So the run has a second half. It appends benchmark-only
[stage-bench-*]sections to the deployed configuration — each with that stage’s own sensible defaults, and every edge pointing at one sharedbench-discardsink (pepsi-stage-discard), so nothing an isolated stage advances can escape into the live pipeline or onto the network — and then injects a sized message straight at each of them with thepepsi.workqueue_injectSQL function, which is exactly what a stage that originates mail does (StageContext::enqueue_new).workqueue_injectnotifies theworkqueuechannel, so the dispatcher picks the row up at once, the stage runs as an ordinary worker, and its cost lands inpepsi.stage_statslike any other. Injecting rather than sending also takespepsi-ingress, SMTP and the admission limits out of the measurement, so an isolated figure is the stage’s own cost and nothing else.isolated stage
program
what is configured, and what it therefore measures
bench-discardpepsi-stage-discardDISPOSITION = success,BOUNCE = no— the shared sink, so its own row is also the pipeline’s bare per-message floorbench-aliasespepsi-stage-aliasesa three-entry
ALIASESmap (a two-target expansion, a transitive hop, an@domaincatch-all) and a recipient that expandsbench-routepepsi-stage-routeone explicit
ROUTEScarve-out plusMANAGED_STAGE; the recipient is at a managed domain, so the managed branch is takenbench-vacationpepsi-stage-vacationVACATION_RANGEScovering today (in the server’s timezone) andSUPPRESS_DAYS = 0, so every message really composes and injects a notice rather than taking the “nobody is away” exitbench-autocrypt-learnpepsi-stage-autocrypt-learndefaults, on a message with no
Autocrypt:header — the pass over a message with nothing to learn, which is what all but a sliver of real inbound mail isbench-vks-confirmpepsi-stage-vks-confirma
VKS_HOSTthe bench sender is not, so the message passes through and no link is followed (no network I/O)bench-auto-paypepsi-stage-auto-paythe wallet configured as the deployed
[stage-pay], on a message that is not a payment demand — again the common inbound case. The paying half needs a live merchant and wallet;12/13drive that.bench-secure-linkpepsi-stage-secure-linkdefaults; terminal. The global
[pepsi-secure-link] NOTIFY_STAGEis pointed at the sink for the run, so timing the stage puts no mail on the wire.bench-lmtppepsi-stage-relay-to-lmtpthe local Dovecot socket and the throwaway mailbox
11-lmtp-test.shprovisions; skipped unless both are presentbench-milterpepsi-stage-milterskipped unless a milter is named in
BENCH_MILTER_ADDR(e.g.inet:8890@127.0.0.1)An isolated stage whose program is not installed, or whose external dependency is missing, is dropped before the sections are written; and if the dispatcher does not come back up with them, the append is rolled back and the isolated half is abandoned rather than leaving the host without its coordinator.
ISO_STAGES=""skips the half entirely.Deliberately not isolated:
pepsi-stage-encrypt,-decrypt,-dot-forward,-srs,-dkim-signand-if— the deployed pipeline routes through all of them, so the scenarios already time them in situ, which is the more honest figure.Reading the matrix. Body-reading stages (arc verify, decrypt, detect-language, encrypt, dkim-sign, secure-link, relay/maildir) climb with size; metadata-only stages (if-branches, srs, whitelist, aliases, route) stay flat. The relay stages (
smarthost,internet,delaytest,bench-lmtp) are network- or MDA-inclusive — they hold the outbound transaction open, so their figure is the next hop’s round-trip, not local CPU; the benchmark prints an explicit note when such a stage ran more times than messages were injected, i.e. when the next hop deferred and the figure averages retries in.``09-latency-bench.sh`` — idle end-to-end latency. Runs the whole inject→observe loop on the ADMIN host (one clock, no ssh-per-poll) for
Nshort messages (default 200) and reports min/mean/median/p95/max, plus the sum of mean per-stage processing time as a floor.POLL_MS(default 2) is how often it looks for the delivered message, and therefore the quantisation of every figure it prints, so it has to be small against the latency being measured: a poll interval near the real latency rounds the distribution into a few buckets and lets the poll rather than the pipeline decide the p95.``10-goodput-bench.sh`` — saturation goodput, three load paths, two instruments. The three paths are A pure pipeline (local inject → local Maildir), B network + front-end (remote sender → local Maildir) and C outbound relay (authorized
ALICEorigin → relay-to-internet →CAROL’s MX). Each is measured twice:a fixed-burst drain (
run_burst) offersBURST_MSGSmessages flat out and times the queue to empty, reporting the rate and the host’s CPU, disk%utiland link utilisation over exactly that window;a ramp (
run_ramp) doubles the offered load everySTEP_SECSuntil the queue is overSAT_QUEUEand still growing.
The burst is the figure; the ramp is the shape. A ramp attributes completions to whichever fixed window they land in, and starting a generator costs an ssh round-trip and a Python start-up on the injecting host — a large share of a five-second window when that host is remote, and larger still when it is small, so the ramp reads well below the burst for the same path. Read the ramp for where the queue starts growing and what the host is doing there.
The burst for each scenario runs before that scenario’s ramp, because a ramp deliberately ends saturated: it leaves a large backlog and a Maildir with as many files in it, and a burst measured on top of that recovery reads far below one on a settled machine.
A burst aimed at a mailbox on another host uses
BURST_MSGS_REMOTE(2 000) rather thanBURST_MSGS(20 000), and empties that mailbox afterwards. Both are about the next script rather than this one: tens of thousands of messages relayed to a small peer leave it delivering for a long time, and every helper that finds a message does it by scanning the mailbox, so a peer still holding them makes the next script’s lookups time out even though it delivered every message.Four things about how it ramps are worth knowing before reading a number out of it.
The ramp is geometric. Each step multiplies the offered load by
RAMP_FACTOR(default 2) instead of adding a fixedRATE_UNIT. A linear ramp cannot serve both ends of the range this suite is pointed at: a small VM and a large server plateau orders of magnitude apart, so a unit that finds the small host’s knee stops far short of the big one and reportsMAX_STEPSas its ceiling. What it costs is resolution at the knee — the offered rate that provoked the plateau is bracketed only to within a factor ofRAMP_FACTOR. The sustained figure is not affected: that is the measured completion rate.RAMP_FACTOR=1gives a linear ramp ofRATE_UNITsteps.Saturation is a growing queue, not a flat rate. The ramp stops when the queue is “over
SAT_QUEUEand still growing forSTALL_STEPS(default 2) consecutive steps”, which is what saturation is — more admitted than completed. A plateau in the processed rate would be the wrong signal: the per-window rate keeps wobbling either way long after the pipeline is flat out, so a plateau test can run out of steps while the queue is already deep.It reports three utilisations, and only one of them may be extrapolated from. CPU, the
%utilof the disk holdingPGDATA, and the NIC’s link utilisation are sampled over each burst and each ramp step. CPU tracks the load; the link is nowhere near binding but could be on a busier one; disk%utilmust not be used, because it counts wall time with a non-empty request queue rather than device saturation, and on a parallel NVMe device it need not rise with throughput at all. The summary’s extrapolated column therefore scales by CPU and caps the result at the best rate any all-local path actually achieved (see Extrapolating past a remote bottleneck).It raises the pipeline’s own worker count, not just the admission caps:
BENCH_STAGE_PAR(default 16) is written asPARALLELISMinto every[stage-*]section for the run. The deployed default of 4 exists so that an out-of-the-box install fits an out-of-the-boxmax_connections; measuring it and calling the answer the machine’s ceiling is measuring the default. Because each stage worker holds exactly one database connection, the script then printsΣ PARALLELISMplus the dispatcher and server pools against PostgreSQL’smax_connectionsand warns when it does not fit — a worker that cannot get a connection reports database-overload, the dispatcher requeues and throttles, and the ramp quietly reports a lower number, so an unchecked budget error looks exactly like a capacity figure.
26.3. Rate-limit reporting¶
The parallel load generator records the SMTP reply of every rejection, so
Pepsi’s own front-end throttle ([pepsi-ingress] CONN_RATE_PER_SECOND /
CONN_RATE_BURST) shows up as 421s and is reported with the knob to
raise.
For the outbound path (scenario C), which touches MTAs Pepsi does not
control, the script reads Pepsi’s view of the next hop — state.bounce SMTP
codes/text, pepsi.tls_session, the relay journal — and, if it sees 4xx /
“rate”/”too many”/”try again”/”throttle” deferrals, reports external
rate-limiting with concrete fixes:
On the remote MTA (Postfix): raise
smtpd_client_connection_rate_limit,smtpd_client_message_rate_limit, the*_connection_count_limit, widenanvil_rate_time_unit, or exempt the Pepsi host’s IP viasmtpd_client_event_limit_exceptions/ a dedicated transport.On the Pepsi side (to stay under a limit you cannot change): lower
[stage-internet] PARALLELISMand pace retries viaRETRY_INITIAL/RETRY_MAX_INTERVAL/RETRY_FACTOR.
26.4. Useful knobs¶
script |
knobs |
|---|---|
all |
|
08 |
|
09 |
|
10 |
|