.. This file is part of PEPSI. Copyright (C) 2026 GNUnet e.V. PEPSI is free software; you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation; either version 3, or (at your option) any later version. PEPSI is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU Affero General Public License for more details. .. _bench-suite: ================== Benchmark Suite ================== The ``tests/NN-*-bench.sh`` scripts measure how fast the live Pepsi pipeline processes mail. They reuse the same hosts, accounts and helpers as the :doc:`correctness suite ` (``tests/test-accounts.ini``, ``tests/lib.sh``) plus a shared ``tests/bench-lib.sh``, and are driven the same way:: tests/08-perstage-bench.sh tests/test-accounts.ini tests/09-latency-bench.sh tests/test-accounts.ini tests/10-goodput-bench.sh tests/test-accounts.ini # or all three in order: make benchmarks ACCOUNTS=tests/test-accounts.ini They are **not** part of ``make integrationtests`` (they are slow and deliberately load the MTA). Each one runs ``preflight_common`` first, so it SKIPs cleanly if the live hosts/DB/services are unreachable. Run ``tests/01-deploy.sh`` first so the pipeline — including the ``DAVE`` local-Maildir delivery target — exists. This chapter describes the method. For how to read the results — per-stage cost, idle latency and saturation goodput — and what to tune, see the :doc:`Performance ` chapter. How the numbers are obtained ============================ * **Per-stage cost** comes from the dispatcher's own timing: ``pepsi-dispatch`` records every stage execution into ``pepsi.stage_stats(stage, messages, duration_us)``. Snapshotting that table before/after a message gives exact microseconds-per-message per stage — for a stage reached by an SMTP-sent message and for one injected straight at it alike, since both run as ordinary dispatched workers. The benchmarks lower ``[pepsi-dispatch] STATS_INTERVAL`` (and ``POLL_INTERVAL``) so the counters flush promptly, and also set ``[pepsi] LOG`` to a quiet level (``warn`` by default, overridable via ``BENCH_LOG_LEVEL``) so per-message logging does not inflate the timings — that knob is read by every component, so it covers the dispatcher *and* the stage workers it spawns (which take no log flag of their own). Both ``pepsi-dispatch`` and ``pepsi-ingress`` are restarted so the new level takes effect, and the benchmarks **restore the deployed config on exit** (a one-shot ``pepsi.conf.bench.bak`` backup on the ADMIN host, replayed by a ``trap … EXIT``). * **Goodput** is the delta of the global ``pepsi.dispatch_stats.messages_processed`` counter over each 5-second window. It is flushed every ``STATS_INTERVAL``, so at a high rate it lags the truth by up to one interval; the ``deliv/s`` column beside it — files appearing in the destination ``Maildir/new`` — is the instantaneous cross-check, and the two disagreeing by a step's worth is that lag, not a lost message. * Delivery is cross-checked by counting files in the target ``Maildir/new``. * **The load generator is several processes, not one.** ``load_burst`` splits a burst across ``LOAD_PROCS`` (default 4) OS processes, each running its share of the threads over its own stride of the message range, and merges their JSON. One Python process is one GIL, and at high rates the interpreter is what stops the number going up — which means a single-process generator stops measuring the system under test and starts measuring itself. To keep the inbound pay-to-send (anti-spam) stage from gating the load, the benchmarks pre-whitelist the sender (``bench_whitelist_sender``), so ``check-whitelist`` sets ``state.spam=false`` and ``anti-spam`` short-circuits. .. note:: **The ramp cannot build a backlog through SMTP, and that is not a flaw in the ramp.** ``pepsi-ingress`` commits each admission synchronously — a ``250`` means the message reached the disk — so on a host where admission is the narrower of the two, the queue stays near empty however hard the generator pushes and the ramp measures the front end. To measure the *pipeline* with a deep queue, insert rows straight into ``pepsi.workqueue`` at ``status = 'pending'`` in one transaction and time the drain; the :doc:`Performance ` chapter shows how, and the difference between the two figures is the interesting part. The three benchmarks ==================== 1. **``08-perstage-bench.sh`` — per-stage cost vs message size.** Prints one ``stage × size`` matrix of µs/msg covering **every stage Pepsi has**, filled in by two halves that answer two different questions: the *scenarios* time the stages of the deployed pipeline **in situ**, and the *isolated stages* time every stage program the deployed pipeline routes no message through. **The scenarios.** A single message only traverses one path, so it can only time the stages on that path; covering the whole bidirectional, branching pipeline therefore takes several messages — one per path. The benchmark drives a set of ``SCENARIOS``, each a message whose journey reveals a different group of stages, and accumulates their per-stage costs into the matrix: .. list-table:: :header-rows: 1 :widths: 18 30 52 * - scenario - message - stages it newly reveals * - ``inbound-local`` - ``CAROL`` → ``DAVE`` (foreign inbound → local Maildir) - ``init`` (arc), ``route`` (if), ``decrypt``, ``detect-language``, ``block-language``, ``check-whitelist``, ``anti-spam`` (short-circuit), ``list`` (list router), ``forward`` (dot-forward), ``local`` (maildir) * - ``inbound-relay`` - ``CAROL`` → ``ALICE`` (inbound → onward relay) - ``srs``, ``dkim-sign-relay``, ``smarthost`` (the non-local recipient is peeled by ``local`` onto that tail) * - ``outbound`` - ``ALICE`` → ``CAROL`` (local-origin submission → direct relay) - ``list-out`` (list router), ``edit-settings``, ``auto-whitelist``, ``delay-route`` (if), ``encrypt``, ``dkim-sign``, ``internet`` (relay-to-internet) * - ``bounce`` - ``CAROL`` → ``dave-broken`` (local delivery fails) - ``bounce`` (the maildir helper fails → a bounce ``SideClone`` routes to the bounce stage → ``dkim-sign``) * - ``bounce-language`` - ``CAROL`` → ``DAVE``, **French** body (blacklisted language) - ``bounce-language`` (``block-language`` bounces an ``fr`` message) * - ``payment`` *(opt-in)* - ``CAROL`` (not whitelisted) → ``DAVE`` (pay-to-send gate) - ``bounce-payment`` (shortens ``PAYMENT_DEADLINE`` for the run) * - ``delay`` *(opt-in)* - ``ALICE`` → ``blackhole.invalid``, ``ENVID=PEPSIDELAY`` (DELAY DSN) - ``delaytest`` (relay-to-smarthost to an unrouteable MTA) The default ``SCENARIOS`` are the five fast ones; append ``payment`` and/or ``delay`` to also time the last two deployed stages (they are slower — merchant- and retry-bound). The ``payment``/``delay`` scenarios and the two opt-in stages need a sender that is not local-origin and the broken local target, all of which ``01-deploy.sh`` provisions. Setting ``SCENARIOS=""`` runs the isolated half alone. **The isolated stages.** ``pepsi-ingress`` always admits a message at ``init`` (``pepsi_common::stage::INITIAL_STAGE``), so a scenario can only ever reach the stages the *deployed* pipeline routes it through. Every stage program that is not wired into that pipeline is therefore unreachable by any message, however constructed. So the run has a second half. It appends benchmark-only ``[stage-bench-*]`` sections to the deployed configuration — each with that stage's own sensible defaults, and every edge pointing at one shared ``bench-discard`` sink (``pepsi-stage-discard``), so nothing an isolated stage advances can escape into the live pipeline or onto the network — and then injects a sized message **straight at each of them** with the ``pepsi.workqueue_inject`` SQL function, which is exactly what a stage that originates mail does (``StageContext::enqueue_new``). ``workqueue_inject`` notifies the ``workqueue`` channel, so the dispatcher picks the row up at once, the stage runs as an ordinary worker, and its cost lands in ``pepsi.stage_stats`` like any other. Injecting rather than sending also takes ``pepsi-ingress``, SMTP and the admission limits out of the measurement, so an isolated figure is the stage's own cost and nothing else. .. list-table:: :header-rows: 1 :widths: 26 30 44 * - isolated stage - program - what is configured, and what it therefore measures * - ``bench-discard`` - ``pepsi-stage-discard`` - ``DISPOSITION = success``, ``BOUNCE = no`` — the shared sink, so its own row is also the pipeline's bare per-message floor * - ``bench-aliases`` - ``pepsi-stage-aliases`` - a three-entry ``ALIASES`` map (a two-target expansion, a transitive hop, an ``@domain`` catch-all) and a recipient that expands * - ``bench-route`` - ``pepsi-stage-route`` - one explicit ``ROUTES`` carve-out plus ``MANAGED_STAGE``; the recipient is at a managed domain, so the managed branch is taken * - ``bench-vacation`` - ``pepsi-stage-vacation`` - ``VACATION_RANGES`` covering today (in the **server's** timezone) and ``SUPPRESS_DAYS = 0``, so every message really composes and injects a notice rather than taking the "nobody is away" exit * - ``bench-autocrypt-learn`` - ``pepsi-stage-autocrypt-learn`` - defaults, on a message with **no** ``Autocrypt:`` header — the pass over a message with nothing to learn, which is what all but a sliver of real inbound mail is * - ``bench-vks-confirm`` - ``pepsi-stage-vks-confirm`` - a ``VKS_HOST`` the bench sender is not, so the message passes through and no link is followed (no network I/O) * - ``bench-auto-pay`` - ``pepsi-stage-auto-pay`` - the wallet configured as the deployed ``[stage-pay]``, on a message that is not a payment demand — again the common inbound case. The paying half needs a live merchant and wallet; ``12``/``13`` drive that. * - ``bench-secure-link`` - ``pepsi-stage-secure-link`` - defaults; terminal. The global ``[pepsi-secure-link] NOTIFY_STAGE`` is pointed at the sink for the run, so timing the stage puts no mail on the wire. * - ``bench-lmtp`` - ``pepsi-stage-relay-to-lmtp`` - the local Dovecot socket and the throwaway mailbox ``11-lmtp-test.sh`` provisions; **skipped** unless both are present * - ``bench-milter`` - ``pepsi-stage-milter`` - **skipped** unless a milter is named in ``BENCH_MILTER_ADDR`` (e.g. ``inet:8890@127.0.0.1``) An isolated stage whose program is not installed, or whose external dependency is missing, is dropped **before** the sections are written; and if the dispatcher does not come back up with them, the append is rolled back and the isolated half is abandoned rather than leaving the host without its coordinator. ``ISO_STAGES=""`` skips the half entirely. Deliberately *not* isolated: ``pepsi-stage-encrypt``, ``-decrypt``, ``-dot-forward``, ``-srs``, ``-dkim-sign`` and ``-if`` — the deployed pipeline routes through all of them, so the scenarios already time them in situ, which is the more honest figure. **Reading the matrix.** Body-reading stages (arc verify, decrypt, detect-language, encrypt, dkim-sign, secure-link, relay/maildir) climb with size; metadata-only stages (if-branches, srs, whitelist, aliases, route) stay flat. The relay stages (``smarthost``, ``internet``, ``delaytest``, ``bench-lmtp``) are **network- or MDA-inclusive** — they hold the outbound transaction open, so their figure is the next hop's round-trip, not local CPU; the benchmark prints an explicit note when such a stage ran more times than messages were injected, i.e. when the next hop deferred and the figure averages retries in. 2. **``09-latency-bench.sh`` — idle end-to-end latency.** Runs the whole inject→observe loop on the ADMIN host (one clock, no ssh-per-poll) for ``N`` short messages (default 200) and reports min/mean/median/p95/max, plus the sum of mean per-stage processing time as a floor. ``POLL_MS`` (default 2) is how often it looks for the delivered message, and therefore the **quantisation of every figure it prints**, so it has to be small against the latency being measured: a poll interval near the real latency rounds the distribution into a few buckets and lets the poll rather than the pipeline decide the p95. 3. **``10-goodput-bench.sh`` — saturation goodput, three load paths, two instruments.** The three paths are **A** pure pipeline (local inject → local Maildir), **B** network + front-end (remote sender → local Maildir) and **C** outbound relay (authorized ``ALICE`` origin → relay-to-internet → ``CAROL``'s MX). Each is measured twice: * a **fixed-burst drain** (``run_burst``) offers ``BURST_MSGS`` messages flat out and times the queue to empty, reporting the rate and the host's CPU, disk ``%util`` and link utilisation over exactly that window; * a **ramp** (``run_ramp``) doubles the offered load every ``STEP_SECS`` until the queue is over ``SAT_QUEUE`` and still growing. **The burst is the figure; the ramp is the shape.** A ramp attributes completions to whichever fixed window they land in, and starting a generator costs an ssh round-trip and a Python start-up on the injecting host — a large share of a five-second window when that host is remote, and larger still when it is small, so the ramp reads well below the burst for the same path. Read the ramp for *where* the queue starts growing and what the host is doing there. The burst for each scenario runs **before** that scenario's ramp, because a ramp deliberately ends saturated: it leaves a large backlog and a Maildir with as many files in it, and a burst measured on top of that recovery reads far below one on a settled machine. A burst aimed at a mailbox on another host uses ``BURST_MSGS_REMOTE`` (2 000) rather than ``BURST_MSGS`` (20 000), and empties that mailbox afterwards. Both are about the *next* script rather than this one: tens of thousands of messages relayed to a small peer leave it delivering for a long time, and every helper that finds a message does it by scanning the mailbox, so a peer still holding them makes the next script's lookups time out even though it delivered every message. Four things about *how* it ramps are worth knowing before reading a number out of it. * **The ramp is geometric.** Each step multiplies the offered load by ``RAMP_FACTOR`` (default 2) instead of adding a fixed ``RATE_UNIT``. A linear ramp cannot serve both ends of the range this suite is pointed at: a small VM and a large server plateau orders of magnitude apart, so a unit that finds the small host's knee stops far short of the big one and reports ``MAX_STEPS`` as its ceiling. What it costs is resolution *at* the knee — the offered rate that provoked the plateau is bracketed only to within a factor of ``RAMP_FACTOR``. The *sustained* figure is not affected: that is the measured completion rate. ``RAMP_FACTOR=1`` gives a linear ramp of ``RATE_UNIT`` steps. * **Saturation is a growing queue, not a flat rate.** The ramp stops when the queue is "over ``SAT_QUEUE`` and still growing for ``STALL_STEPS`` (default 2) consecutive steps", which is what saturation *is* — more admitted than completed. A plateau in the processed rate would be the wrong signal: the per-window rate keeps wobbling either way long after the pipeline is flat out, so a plateau test can run out of steps while the queue is already deep. * **It reports three utilisations, and only one of them may be extrapolated from.** CPU, the ``%util`` of the disk holding ``PGDATA``, and the NIC's link utilisation are sampled over each burst and each ramp step. CPU tracks the load; the link is nowhere near binding but could be on a busier one; disk ``%util`` **must not be used**, because it counts wall time with a non-empty request queue rather than device saturation, and on a parallel NVMe device it need not rise with throughput at all. The summary's extrapolated column therefore scales by CPU and caps the result at the best rate any all-local path actually achieved (see :ref:`perf-extrapolation`). * **It raises the pipeline's own worker count**, not just the admission caps: ``BENCH_STAGE_PAR`` (default 16) is written as ``PARALLELISM`` into every ``[stage-*]`` section for the run. The deployed default of 4 exists so that an out-of-the-box install fits an out-of-the-box ``max_connections``; measuring it and calling the answer the machine's ceiling is measuring the default. Because each stage worker holds exactly one database connection, the script then prints ``Σ PARALLELISM`` plus the dispatcher and server pools against PostgreSQL's ``max_connections`` and **warns when it does not fit** — a worker that cannot get a connection reports database-overload, the dispatcher requeues and throttles, and the ramp quietly reports a lower number, so an unchecked budget error looks exactly like a capacity figure. Rate-limit reporting ==================== The parallel load generator records the SMTP reply of every rejection, so Pepsi's **own** front-end throttle (``[pepsi-ingress] CONN_RATE_PER_SECOND`` / ``CONN_RATE_BURST``) shows up as ``421``\ s and is reported with the knob to raise. For the outbound path (scenario C), which touches MTAs Pepsi does **not** control, the script reads Pepsi's view of the next hop — ``state.bounce`` SMTP codes/text, ``pepsi.tls_session``, the relay journal — and, if it sees ``4xx`` / "rate"/"too many"/"try again"/"throttle" deferrals, reports **external rate-limiting** with concrete fixes: * On the remote MTA (Postfix): raise ``smtpd_client_connection_rate_limit``, ``smtpd_client_message_rate_limit``, the ``*_connection_count_limit``, widen ``anvil_rate_time_unit``, or exempt the Pepsi host's IP via ``smtpd_client_event_limit_exceptions`` / a dedicated transport. * On the Pepsi side (to stay under a limit you cannot change): lower ``[stage-internet] PARALLELISM`` and pace retries via ``RETRY_INITIAL`` / ``RETRY_MAX_INTERVAL`` / ``RETRY_FACTOR``. Useful knobs ============ .. list-table:: :header-rows: 1 :widths: 12 88 * - script - knobs * - all - ``BENCH_STATS_INTERVAL``, ``BENCH_POLL_INTERVAL``, ``BENCH_LOG_LEVEL``, ``MSGS_PER_CONN``, ``LOAD_PROCS``, ``BENCH_WHITELIST_NAME``, ``BENCH_CONFIRM_GRACE``; and, for the runs that call ``bench_set_admission``, ``BENCH_CONN_RATE`` / ``BENCH_CONN_BURST`` / ``BENCH_MAX_CONN`` / ``BENCH_DB_POOL`` / ``BENCH_RELAY_PAR_MULT`` / ``BENCH_QUEUE_LIMIT`` / ``BENCH_STAGE_PAR`` / ``BENCH_DISPATCH_POOL`` * - 08 - ``SIZES``, ``REPEAT``, ``SCENARIOS``, ``PERMSG_TIMEOUT``, ``BOUNCE_TIMEOUT_S``, ``PAY_DEADLINE_S``, ``PAY_TIMEOUT_S``; and for the isolated half ``ISO_STAGES``, ``ISO_SIZES``, ``ISO_REPEAT``, ``ISO_TIMEOUT``, ``ISO_LOCAL_RCPT``, ``ISO_SENDER``, ``ISO_LMTP_SOCKET``, ``ISO_LMTP_USER``, ``BENCH_MILTER_ADDR`` * - 09 - ``N``, ``WARMUP``, ``MSG_BYTES``, ``POLL_MS``, ``LATENCY_MAX_MS``, ``PERMSG_TIMEOUT_MS`` * - 10 - ``RATE_UNIT``, ``RAMP_FACTOR``, ``PAR_UNIT``, ``PAR_MAX``, ``MAX_STEPS``, ``STEP_SECS``, ``MSG_BYTES``, ``STALL_STEPS``, ``SAT_QUEUE``, ``DRAIN_TIMEOUT``, ``WARMUP_MSGS``, ``WARMUP_DRAIN``, ``RUN_A``/``RUN_B``/``RUN_C``; and for the burst half ``BURST_MSGS``, ``BURST_MSGS_REMOTE``, ``BURST_PAR``, ``BURST_TIMEOUT``, ``BURST_SETTLE``