27. Performance

What the Pepsi pipeline spends its time on, how to read the figures the benchmark scripts under tests/ produce (08-perstage-bench.sh, 09-latency-bench.sh and 10-goodput-bench.sh), and which knobs move them. How the scripts obtain their figures is in the Benchmark Suite chapter.

Note

Apart from one dated run under Reference measurements, this chapter quotes no figures. Absolute numbers depend on the host, the database settings and the configured pipeline far more than on anything in Pepsi, so run the suite on your own hardware. What is portable is the shape of the results — which stages are cheap, which scale with message size, and what kind of thing the throughput ceiling turns out to be — and that is what this chapter describes.

Run the benchmarks at a quiet log level: the scripts set [pepsi] LOG (every component honours it — the dispatcher and the stage workers it spawns) to warn for the duration, so per-message logging does not inflate the figures.

27.1. How the pipeline spends its time

Five properties of the architecture set the numbers. The last three follow from one principle: at mail rates the pipeline’s cost is dominated by how often it makes the database coordinate rather than by the work a stage does, so it coordinates only where the information genuinely is.

  • Every stage pays a small fixed per-message database cost. A stage worker loads its row in one SELECT (the row and its applicable per-address settings, in a single round-trip) and commits its outcome in one terminal UPDATE/DELETE (or one fan-out stored function). That pair of round-trips, plus worker scheduling, is the floor every message pays at every hop, independent of message size.

  • Several mechanisms remove or amortise those round-trips. The dispatcher claims work for all stages in a single call (per stage, the fair claim reads at most 64 waiting senders by a loose index scan and at most the stage’s capacity of each one’s oldest rows, so a deep backlog does not make it slower; with FAIR_SCHEDULING = no a LATERAL … LIMIT stops each stage’s index scan at its capacity) and pipelines up to QUEUE_LIMIT messages to each worker, so completing a message needs no prior coordinator round-trip. The load-and-advance hot path runs under READ COMMITTED rather than SERIALIZABLE — safe because at most one writer ever touches a given row (the dispatcher is the sole claimer). A body-free stage’s worker loads and commits a whole pipelined batch of messages in one SELECT … = ANY and one arrayed UPDATE. And stage fusion ([pepsi] ALLOW_FUSION, on by default) runs a fast body-free successor whose FUSION is on — by default if, srs, list, check-whitelist, auto-whitelist, block-language, discard and edit-settings (that last one for non-control messages only, which it can decide from metadata) — inside its predecessor’s worker process on the already-loaded row, collapsing a run of n such stages from n loads + n advances + n hand-offs into a single load, the stages’ work, and one terminal write. Fusion changes no outcome: a stage that cannot be fused (it needs the body, is not folded into the unified binary, or its FUSION is off) simply falls back to a normal dispatched hop. Fusion needs the unified pepsi binary, so it is inert in a multibin development build.

  • Hand-offs between stages are deliberately silent. The dispatcher learns there is work from a LISTEN on the workqueue channel, but a stage advancing a message does not notify. It does not need to: the worker reports each finished message on its standard output, and the coordinator answers that report with the same full capacity-bounded claim a notification would have triggered. A notification on the advance path would tell the dispatcher something it is about to be told anyway, over a cheaper channel.

    Notifications are far from free: PostgreSQL holds a database-wide AccessExclusive lock from the moment a transaction queues one until it commits, so notifying transactions cannot group-commit and each hop would serialise the whole database’s commit path. A notification is therefore issued only where the information is genuinely new — by writers outside the dispatcher’s loop: message admission (coalesced, once per burst rather than once per message), workqueue_inject, workqueue_resume, pepsi-keydisc(1) releasing parked mail, pepsi-failure-bouncer(1) and the pepsi-queue(1) repair commands.

    POLL_INTERVAL remains the safety net behind all of it: a wake-up that is never sent, or is lost, costs latency and never a message.

  • Stage transitions do not wait for the disk. A stage worker’s connection runs with synchronous_commit off ([pepsi-postgres] WORKER_SYNCHRONOUS_COMMIT, default off), so the pipeline’s writes group-commit instead of each paying its own flush. Message admission is unaffected — ingress always commits synchronously, so a 250 still means the message reached the disk. The exposure this adds is that a machine crash can lose the last few hundred milliseconds of stage transitions, whereupon the message is found at the previous stage and runs through it again — which is exactly what already happens when a worker is killed between doing its work and committing.

  • Bookkeeping has exactly one writer. The per-stage and global counters are accumulated in the dispatcher’s memory and written in a single transaction — every STATS_INTERVAL, when the pipeline goes idle, and on shutdown. No worker writes them: stage_stats holds one row per stage, so a counter written per message would put every worker of a stage in a queue behind that stage’s single row — bookkeeping contending where the mail itself never does, since exactly one writer ever touches a given message. A stage fused into another stage’s pass is therefore reported to the dispatcher on the status line that pass already writes, and folded into the counters it already keeps, so per-stage accounting stays exact with no write at all. The cost is the usual one for statistics: deltas not yet flushed are lost if the dispatcher dies.

27.2. Reading the per-stage matrix

The dispatcher times every stage execution, so 08-perstage-bench.sh measures the cost of a stage directly, with no instrumentation of the stage code: the wall time a single message spends inside each stage (worker reads the row, runs the stage logic, commits the outcome), in milliseconds per message, for each message size swept. So that each stage is observable as its own dispatched worker, this benchmark runs with [pepsi] ALLOW_FUSION = no; the latency and goodput benchmarks leave fusion on, measuring the system as deployed.

The matrix has two halves (see the Benchmark Suite chapter for both lists): stages timed in situ, by sending one message down each path of the deployed pipeline, and stage programs the deployment routes no message through, timed in isolation by injecting a message straight at a benchmark-only section. The isolated figures carry no pepsi-ingress or SMTP cost at all.

Three classes of row emerge, and one measurement artefact worth naming.

  • Metadata-only stages are flat and cheap. if (the branchers), srs, list, aliases, route and discard declare Load::Metadata — they never touch the message body — and cost the same at every size. check-whitelist, block-language and vacation behave the same way in practice: they declare Load::Headers, so they pull the header block but never the body, and the header block does not grow with the message.

    Such a stage’s isolated figure — one stage, a warm worker, nothing else in the pipeline — is the cost of the hop itself: the single load SELECT, the single terminal UPDATE, and the worker’s own bookkeeping. Its in-situ figure is the same hop with the rest of the chain’s workers alive around it, and is higher. So “the per-hop floor” is not one number: it is the hop’s own work plus being one of many. That is what stage fusion removes, and it is why the end-to-end latency measured with fusion on is far less than the sum of the matrix’s column.

  • Body-reading stages scale with message size. detect-language is the most expensive single operation in the pipeline and essentially linear in the body, followed by decrypt and secure-link; a size-based if in front of detect-language is the single change that saves the most on large messages. dkim-sign, local — the setuid Maildir writer handling the full message — and init (ARC verify, hash and seal) climb much more gently.

  • The stages that only *inspect* sit at the floor. autocrypt-learn, vks-confirm and auto-pay all declare Load::Full, but for ordinary mail they only look, so their cost rises only with the cost of loading a larger body. Adding them to a pipeline costs about one more hop, not one more scan.

  • The relay rows are not comparable to any of it. smarthost, internet, delaytest and bench-lmtp hold an outbound SMTP or LMTP transaction open, so their figure is another host’s round-trip and that host’s own processing. It need not even rise monotonically with size, and it says nothing about the Pepsi host — see Figures that belong to another host.

  • The artefact: cold workers. The sweep moves from size to size more slowly than [pepsi-dispatch] WORKER_IDLE_TIMEOUT (default 5 s), so a benchmark-only stage’s worker can be reaped between sizes and the next cell pays to have one exec’d. The tell is a jump of the same few milliseconds whatever the message size, alternating with the warm value along a row. The warm value is the floor; the cold one is the floor plus a process start.

anti-spam appears in the matrix cleared: the benchmarks whitelist the sender, so it short-circuits on state.spam and never looks at the body. Its gating cost is whatever the Taler merchant backend’s latency is, which is a property of the merchant rather than of Pepsi. Whitelisting is what keeps it off the common path.

27.3. Figures that belong to another host

Some figures a benchmark run prints are measurements of the Pepsi host talking to another machine, and are bounded by that machine:

  • The relay stages’ per-stage cost, as above: a remote round-trip plus remote processing, not what it costs Pepsi to relay a message.

  • Outbound goodput (scenario C of the goodput benchmark). Relaying is limited by how fast the receiving MX accepts. If that host is itself a Pepsi deployment, its [pepsi-ingress] CONN_RATE_PER_SECOND/CONN_RATE_BURST default to one new connection a second per source, and a relay opening a connection per message is answered 421 and defers — nothing is lost, and nothing is measured either. Raise them on the receiving side first. Once they are raised, what remains is the receiving host’s own speed; a small peer accepts a burst quickly and then spends a long time delivering it.

  • ``n/a`` cells at large sizes on the smarthost path. Postfix’s default message_size_limit is 10 240 000 bytes, and a Postfix smarthost refuses larger messages with 552. The sweep records the refusal and skips the larger sizes on that path rather than reproducing it.

27.3.1. Extrapolating past a remote bottleneck

The goodput summary prints an extrapolated column for exactly this case.

Warning

An extrapolation is arithmetic performed on measurements, not a measurement, and it is an upper bound: it assumes the one resource it scales stays the binding one, which is exactly what stops being true as you approach it. Quote the measured figures.

The method is the obvious one: if a run used a fraction u of some local resource and something else was the limit, then removing that something else would let the rate rise by up to 1/u. Which resource decides whether the answer means anything:

  • Disk ``%util`` must not be used. iostat’s %util is the share of wall time the request queue was non-empty; on an NVMe device that serves dozens of requests in parallel it says “something was outstanding”, not “the device is full”, and it does not rise with throughput in a way that can be scaled by.

  • Network is reported because it would be the first to become binding on a busier link, but is rarely close.

  • CPU tracks the load and is what the projection uses.

And the projection is capped at the best rate any all-local path (no remote host in it) achieved in the same run: you cannot relay messages faster than you can process them. What the extrapolation licenses is the shape of the claim — the outbound path is limited by the peer, by something like this factor — not the digit. The honest way to close the gap is a second machine of the same class on the other end.

27.4. End-to-end latency (idle)

With the system otherwise idle, 09-latency-bench.sh injects short messages one at a time and polls for each one’s arrival in the destination Maildir, timing the whole inject→deliver round-trip across the inbound path, and reports min/mean/median/p95/max. Read the median: a single slow sample — a worker being forked, a checkpoint — is enough to pull the mean above the p95.

The script also prints the sum of the mean per-stage processing times over the same run. On an idle system the end-to-end latency is close to that sum: dispatcher scheduling and the hand-off between stages add little per hop, and the advance path is driven by the worker’s status report rather than by a notification, so POLL_INTERVAL does not gate it.

Two properties of the harness keep the figure honest. The sampler reads each file in the destination Maildir/new once per size it is seen at, rather than every file on every poll, so a mailbox left full by an earlier run does not turn the measurement into one of the harness’s own glob. And it opens a connection per sample, which at low latency is faster than a deployed CONN_RATE_PER_SECOND allows, so it raises the admission caps for the run and reports a refusal at submission separately from a message that was accepted and never arrived.

Idle latency is the per-message work a single message pays on an empty queue. It is not the reciprocal of the saturation throughput below — the two are limited by different things.

27.5. Saturation goodput

10-goodput-bench.sh measures each of three load-and-delivery paths twice, and the two instruments are not equally good.

A ramp multiplies the offered load every STEP_SECS until the queue is over SAT_QUEUE and still growing. A fixed-burst drain offers BURST_MSGS messages flat out and times the queue to empty. Both raise, for the duration, the admission caps ([pepsi-ingress] CONN_RATE_PER_SECOND/CONN_RATE_BURST/ MAX_CONNECTIONS and its DB_POOL_SIZE) and PARALLELISM on every [stage-*] section, and both restore the deployed configuration on exit. Both sample the host’s CPU, the %util of the disk under PGDATA and the NIC’s link utilisation.

Quote the burst. A ramp attributes completions to whichever fixed window they land in, and getting a generator going costs an ssh round-trip and a Python start-up on the injecting host — a large fraction of the window when that host is remote, and larger still when it is small — so a ramp reads well below the burst for the same path. The ramp is worth reading for the shape of the knee — where the queue starts growing, what the host is doing — but not for a rate.

The paths, and what limits each:

path

what limits it

A local inject → local Maildir

often its own load generator, which runs on the host under test; use it to confirm that the network adds nothing, not as a figure to quote

B foreign sender → local Maildir

this host. The inbound figure to quote

C local origin → relay to a foreign MX

the receiving host (see Figures that belong to another host)

pipeline only, SQL-fed drain (see Reproducing)

this host, with SMTP admission taken out of the path entirely

Reading them together:

  • A against B isolates the network. The two differ only in where the sender is; per-stage costs that match to within noise mean inbound authentication of foreign mail is not a per-message cost worth seeing.

  • The SQL-fed drain against B isolates admission. The difference is pepsi-ingress accepting connections and committing each admission synchronously — a 250 means the message reached the disk — and it is the price of that guarantee, not a defect.

  • Check that nothing inside the database contends. Sample pg_stat_activity and pg_locks through a drain. Backends sitting in Client/ClientRead are PostgreSQL waiting for the workers, not the other way round; an ungranted lock or IO/WALSync waits say the database is the limit. The ramp’s fail/s and serr/s columns must read zero at every load level: a valid message must never fail and the single-claimer queue must never serialise, so a non-zero value is a correctness defect rather than a capacity limit.

  • Past the knee, throughput can fall rather than flatten. A Pepsi under more load than it can take loses nothing, but it does not necessarily degrade gracefully to its maximum rate; the admission caps are what keep a burst from provoking this, and they are the knob to reach for.

Note

A saturated queue only ever reports its outermost constraint, so a throughput figure on its own says nothing about what is limiting it. The checks that settle it are pg_locks and pg_stat_activity under load, the host’s own CPU/disk/network, and — as path C shows — what the next hop is doing. Quote a figure alongside that evidence, and re-run the same checks after a change.

A message costs about two database transactions per dispatched hop — the load and the terminal write — plus admission, which is the quantity the coordination properties above are really about, and the one to watch when adding a stage. The headroom lies there: fewer stages on the path, and stage fusion, which collapses a run of fusible body-free stages into one load, one terminal write and one round of bookkeeping instead of n.

27.6. Tuning notes

  • Parallelism and the connection budget. Each stage worker is a process that holds one database connection, so the live connection count is approximately Σ PARALLELISM over busy stages, plus the dispatcher pool and the pepsi-ingress/pepsi-httpd pools. This sum must stay under PostgreSQL’s max_connections. The default PARALLELISM is 4 and the default DB_POOL_SIZE is 1 for pepsi-ingress and pepsi-httpd (the dispatcher’s own pool defaults to 8, pepsi-telemetry’s to 2) precisely so a default install fits a default server. A pipeline with many stage sections can still exceed Debian’s default max_connections = 100 if every stage is busy at once. Raising PARALLELISM to push throughput, or DB_POOL_SIZE to admit more concurrent sessions, means re-checking that budget. If it is exceeded, workers that cannot get a connection report the database-overload status and the dispatcher requeues the message and throttles the busiest stage rather than losing mail — see pepsi-dispatch (Database-overload backpressure). That does not produce an error, only a lower throughput, which is why 10-goodput-bench.sh prints the sum against max_connections before it measures anything. The remedy for sustained pressure is to raise max_connections or lower PARALLELISM, not to rely on the throttle.

  • Fewer hops always helps; more workers helps only when the worker count is what is in the way. A dispatched hop costs a round-trip pair — the worker’s load SELECT and its terminal UPDATE — plus the scheduling around it, so the goodput ceiling moves with path length whatever else is true. (It is not a statistics write: no worker writes stage_stats at all, see Bookkeeping has exactly one writer above.) Shortening the path, or fusing the hop away, is therefore the change that always pays.

    PARALLELISM is the one that depends on the machine. On a small host the knee is CPU or the write-ahead log, and more workers only means more of them waiting. On a large one, a few workers per stage, each blocked on its own database round-trip, can leave the machine nearly idle with no lock contention anywhere, and raising PARALLELISM is then the whole story. The rule is the diagnosis rather than a number: if the host is idle and pg_locks is clean at the knee, the worker count is the ceiling and raising it will work; if the host is busy, or the waits are IO/WALSync, it will not. Each worker is one database connection, so how far it can go is a property of max_connections. (Disk %util does not settle the question — see Extrapolating past a remote bottleneck.)

  • Leave stage fusion on. [pepsi] ALLOW_FUSION (on by default) runs a fusible body-free successor inside its predecessor’s worker on the already-loaded row, so a run of n such stages costs one load, one terminal write and one round of bookkeeping instead of n. The only reason to turn it off is to measure stages separately, as the per-stage benchmark does. Setting FUSION = yes on a cheap metadata-only stage that the fusion registry supports is worth more than any pool knob.

  • Pipeline depth before parallelism. For the same reason, reach for a stage’s QUEUE_LIMIT (messages pipelined to one worker) before its PARALLELISM. QUEUE_LIMIT lifts a stage’s in-flight capacity to QUEUE_LIMIT × PARALLELISM and removes a coordinator round-trip per message without adding worker processes or database connections, so it does not enter the budget above; and a body-free stage’s worker commits a whole pipelined batch in one arrayed UPDATE, which is one round-trip for the whole batch. It defaults to 4 — pipelining is on out of the box — so the knob is often about lowering it: set QUEUE_LIMIT = 1 for stages whose per-message work is long and uneven, chiefly the network relays, where a pipelined message would otherwise wait behind a slow sibling.

  • Drop work you do not need. The cheapest stage is one that is not in the chain — cheapest twice over, since it costs neither its own work nor a hop’s share of the coordination. detect-language in particular is worth omitting, or guarding with a size branch, unless language filtering is actually used.

  • Leave ``WORKER_SYNCHRONOUS_COMMIT`` off unless a stage in your pipeline has an external side effect that must not be repeated even across a power loss. Turning it on costs roughly a write-ahead-log flush per stage hop. What that is worth depends entirely on the disk — considerable on a SATA SSD, small on a fast NVMe device — so measure it where you are before trading it away.

  • Front-end caps are separate from pipeline throughput, and they bind first. [pepsi-ingress] CONN_RATE_PER_SECOND / CONN_RATE_BURST / MAX_CONNECTIONS are a per-source admission policy, and the defaults are deliberately strict — CONN_RATE_PER_SECOND and CONN_RATE_BURST both default to 1, one new connection a second from a given address. That is the right default for a host on the open internet and completely wrong for a host another of your own machines relays through, and it is the first thing to check when a neighbour of yours seems slow to accept mail. Raise them on the receiving side, per source if you can. The benchmarks raise them on the host under test for the duration precisely so they measure the pipeline rather than this.

  • PostgreSQL itself. The benchmarks need no database tuning beyond max_connections, and the suite deliberately leaves the rest at the distribution’s defaults; a figure measured that way is a floor that a tuned database (shared_buffers in particular) would raise, not a ceiling.

27.7. Reference measurements

One run of the three benchmarks, taken on 2026-09-26 on commit 0c6a40ea with the pipeline tests/01-deploy.sh deploys and every knob at the default the Benchmark Suite chapter lists. They describe this host and this commit only; read them for the shape set out in the sections above, not as a figure to expect elsewhere.

The host under test and its peers

item

value

host under test

the ADMIN host, bare metal

CPU

2 x AMD EPYC 9555, 128 cores / 256 threads

memory

256 GiB

storage

PGDATA and the Maildirs on ext4 on one NVMe SSD (Samsung 9100 PRO)

operating system

Debian GNU/Linux 13 (trixie), kernel 6.12

database

PostgreSQL 17.11 on the same host, distribution defaults (shared_buffers 128 MB, synchronous_commit on) except max_connections = 700

set by the suite for the run

STATS_INTERVAL 2 s; POLL_INTERVAL 2 s (5 s for the goodput run); [pepsi] LOG = warn; ALLOW_FUSION = no for the per-stage run only; the admission caps raised to CONN_RATE_PER_SECOND = 5000, CONN_RATE_BURST = 10000, MAX_CONNECTIONS = 8000, ingress DB_POOL_SIZE = 16; for the goodput run also PARALLELISM = 16 on every [stage-*] section and dispatcher DB_POOL_SIZE = 16 (about 561 of the 700 connections)

ALICE/BOB host

a 2-vCPU, 4 GiB virtual machine running Postfix; the smarthost of the inbound relay path and the sender of path C

CAROL host

a 4-thread Core i3-4005U with 16 GiB running Pepsi itself; the foreign sender of the inbound scenarios and path B, and the receiving MX of the internet stage and path C

Per-stage cost (08-perstage-bench.sh), milliseconds per message, mean of three messages per cell. The bench-* rows are the isolated half, every other row was timed in situ:

stage

1K

10K

100K

1M

5M

10M

20M

init

17.9

17.6

18.0

19.4

27.0

35.4

63.6

route

6.8

6.5

7.0

6.5

7.1

6.6

11.3

decrypt

7.1

7.1

8.4

21.8

76.2

144

286

detect-language

11.9

16.0

23.9

57.0

208

385

741

block-language

4.9

5.0

5.0

5.1

5.4

5.3

9.4

check-whitelist

5.6

5.4

5.9

4.7

4.7

5.3

8.5

anti-spam

4.8

5.0

4.8

6.0

7.0

9.7

19.0

list

6.1

6.2

6.1

6.5

6.4

5.7

10.3

forward

4.3

6.0

6.1

7.6

26.2

11.3

20.2

local

4.6

10.9

11.6

13.4

21.1

27.5

48.9

srs

5.8

6.3

5.8

6.3

14.5

n/a

n/a

dkim-sign-relay

13.0

13.9

13.3

16.6

49.3

n/a

n/a

smarthost

294

267

258

217

347

n/a

n/a

list-out

11.9

6.6

5.8

6.9

10.7

15.1

14.3

edit-settings

9.7

4.8

5.6

5.7

12.8

18.8

25.1

auto-whitelist

10.5

7.0

6.4

6.8

10.3

15.6

15.2

delay-route

10.7

6.1

5.4

5.5

9.2

13.4

15.1

encrypt

10.1

6.9

6.4

11.1

30.8

54.4

91.4

dkim-sign

14.0

13.8

13.0

17.1

40.1

66.9

98.5

internet

245

112

132

256

628

1145

2121

bounce

5.3

5.1

5.0

5.3

5.1

5.2

5.5

bounce-language

5.1

4.8

4.6

5.5

5.2

5.0

5.0

bench-discard

6.5

1.8

1.7

1.7

1.9

6.5

6.6

bench-aliases

7.2

2.0

1.9

1.8

1.9

6.2

6.9

bench-route

6.8

2.0

1.8

1.8

1.7

6.3

6.9

bench-vacation

6.3

1.4

1.6

1.2

1.1

5.9

6.2

bench-autocrypt-learn

5.5

0.7

0.9

1.8

4.5

11.2

17.5

bench-vks-confirm

5.0

0.8

0.7

1.4

3.8

11.0

16.2

bench-auto-pay

5.2

0.8

0.8

1.5

3.6

10.8

16.0

bench-secure-link

24.9

19.3

20.3

27.5

55.6

95.7

175

n/a is the smarthost’s 552 for a 10 MB message, see Figures that belong to another host. The smarthost and internet rows are the round-trip to and the processing on the ALICE/BOB host and the CAROL host respectively, not a cost of this host. The isolated rows show the cold-worker artefact at 1K, 10M and 20M: their warm value, about 2 ms for a metadata-only stage, is the floor.

Idle latency (09-latency-bench.sh), 200 messages of 1 KiB from the local SMTP port to a local Maildir, polled every 2 ms:

delivered

min

median

mean

p95

max

sum of mean stage times

200 of 200

26 ms

28 ms

32 ms

29 ms

448 ms

15.8 ms (10 stages)

Saturation goodput (10-goodput-bench.sh), 2 KiB messages. The burst is the figure; the ramp column is the last step that kept up, and the utilisations are this host’s over the burst:

path

burst (msgs/s)

ramp (msgs/s)

CPU

disk %util

link

A local inject -> local Maildir, 20 000 messages

1162

256

8 %

63 %

0.0 %

B CAROL host -> local Maildir, 20 000 messages

1210

334

7 %

65 %

2.8 %

C ALICE/BOB host -> relay to CAROL host, 2 000 messages

88

109

1 %

11 %

1.8 %

No run recorded a failed message, a serialisation error or an SMTP refusal. Path B is the inbound figure to quote. It ran with the host’s CPU 7 % busy, the situation Tuning notes describes under PARALLELISM; this run did not establish which limit was binding. Path C is the CAROL host’s speed: it had accepted all 2 000 messages when the queue here emptied, and had delivered 296 of them.

27.8. Reproducing

The benchmarks are not part of make integrationtests (they are slow and deliberately load the MTA). Run them against a deployed pipeline with:

tests/08-perstage-bench.sh tests/test-accounts.ini
tests/09-latency-bench.sh  tests/test-accounts.ini
tests/10-goodput-bench.sh  tests/test-accounts.ini
# or all three in order:
make benchmarks ACCOUNTS=tests/test-accounts.ini

Each script lowers the dispatcher’s STATS_INTERVAL/POLL_INTERVAL for the duration of the run so the counters flush promptly, and restores the deployed configuration on exit (including the admission caps the goodput run raises). See the Benchmark Suite chapter for the full methodology and the available tunables (message sizes, sample counts, ramp parameters, admission overrides).

The SQL-fed drain is not one of the scripts, because it deliberately bypasses the part of Pepsi the scripts go through. With the same pacing and PARALLELISM the goodput run sets, insert the backlog directly and time the queue to empty:

INSERT INTO pepsi.workqueue
  (token, mail_from, rcpt_to, from_header, subject, headers, body,
   stage, status, state)
SELECT 'bulk-' || g, 'sender@example.org', ARRAY['dave@<your-host>'],
       'sender@example.org', 'bulk ' || g,
       convert_to('From: sender@example.org' || E'\r\n' || '…' || E'\r\n', 'UTF8'),
       convert_to(repeat('filler' || E'\r\n', 300), 'UTF8'),
       'init', 'pending', '{"local_origin": false, "spam": false}'::jsonb
FROM generate_series(1, 50000) g;

One transaction, so the dispatcher sees the whole backlog at once; then poll SELECT count(*) FROM pepsi.workqueue until it reaches zero. Nothing notifies the dispatcher, so this also exercises the POLL_INTERVAL heartbeat.