25. Test Suite

Pepsi ships an end-to-end test suite under tests/ that exercises a live deployment over SMTP, asserting the observable behaviour of the whole pipeline rather than any single function in isolation. It is deliberately a black-box suite: messages are sent with the ordinary mail command or by raw SMTP, and the assertions read the recipients’ mailboxes and the pepsi.workqueue table.

This complements — it does not replace — the in-tree unit tests (cargo test --workspace), which cover the pure logic (SRS, MIME downgrade, DSN construction, language scoring, DNS classification, …) without a network.

Two things are covered in between, by integration tests that need neither a network nor a database. The migration suites (pepsi-setup/tests/import_*.rs and tests/cli_import.rs) run each MTA importer against a fixture /etc tree and then run the real pepsi-setup binary end to end, up to a --wizard --import that must produce a configuration passing Pepsi’s own validation; see the extending chapter for their layout. The dispatcher lifecycle suite (pepsi-test-stages) needs a PostgreSQL and is run by make check.

25.1. make check: the no-remote-accounts gate

make check is the entry point for everything that needs no live mail accounts. In order, it runs:

  1. make check-telemetry-socket and make check-systemd-units — tests/check-telemetry-socket.sh and tests/check-systemd-units.sh, consistency checks over the systemd units and Debian packaging that need nothing installed and never skip.

  2. cargo test --workspace --exclude pepsi-test-stages --locked — every database-free test in the tree. (Plain cargo test --workspace would include pepsi-test-stages and so need a database; prefer make check, which gates that suite instead of failing on it.)

  3. tests/check-lifecycle.sh — the dispatcher-lifecycle suite, best effort.

  4. tests/maildir-writer-suid.sh, tests/dot-forward-suid.sh and tests/whitelist-suid.sh — the setuid/setgid helper contracts, best effort.

  5. tests/interop-auth.sh — a cross-implementation DKIM/ARC check, best effort. Pepsi’s own round-trip tests sign and verify with the same mail-auth library, so they cannot catch a transform that is consistently wrong in it; this script grades Pepsi-signed messages (all four c= combinations, plus a tampered copy that must be rejected, plus an ARC seal) with two independent implementations, with no DNS involved. dkimpy gets the public keys through its dnsfunc (tests/interop-verify.py) and grades DKIM and the ARC chain; OpenDKIM reads them from a TestPublicKeys file in the opendkim filter’s -t test mode and grades DKIM (it has no ARC verifier; opendkim-testmsg cannot be used, because it verifies only against live DNS). Every DKIM-Signature is graded, the RSA and the Ed25519 one — dkimpy needs PyNaCl for the latter. Each verifier SKIPs its own checks when absent, and the script SKIPs as a whole with neither. On Debian: apt install python3-dkim python3-nacl opendkim (the filter lives in /usr/sbin; set OPENDKIM if it is elsewhere, and disable the opendkim service the package enables — the test never talks to it). PEPSI_INTEROP_REQUIRE=1 turns a missing verifier into a failure.

  6. cargo deny check — advisories, bans, licences and sources. This one is not best-effort: it is skipped only when cargo-deny is genuinely not installed, and a real finding fails the target. Read deny.toml before trusting the gate: its ignore list is not empty.

  7. make docs-linkcheck — every internal cross-reference of the manual must resolve, in every language (a Sphinx dummy build with warnings as errors). Skipped only when sphinx-build is not installed.

The Thunderbird S/MIME interoperability gate is deliberately not in make check — it starts a real (headless, Marionette-driven) Thunderbird and costs about fifteen seconds. It is make interop-thunderbird, and make integrationtests depends on it.

25.1.1. Continuous integration

Several of the steps above skip when a tool or a database is missing — the right default on a developer’s machine, and exactly wrong for an automated run, where a skipped gate reported green looks like a passed one. contrib/ci/ holds jobs in the layout GNU Taler’s CI worker runs (the same one vendor/taler-rust/contrib/ci/ uses): contrib/ci/ci.sh <job> builds the image from contrib/ci/Containerfile with podman and runs contrib/ci/jobs/<job>/job.sh in it with the source tree mounted; contrib/ci/run-all-jobs.sh runs every job in order. The jobs are:

1-gates

The compile-time gates: RUSTFLAGS="-D warnings" cargo build, cargo clippy … -D warnings and cargo fmt --check.

2-check

make check with every best-effort gate armed. The image carries dkimpy, PyNaCl, OpenDKIM, cargo-deny and Sphinx, the job starts PostgreSQL, sets PEPSI_INTEROP_REQUIRE=1, and fails if any line of the transcript reports a SKIP. make check runs as an unprivileged user on a private copy of the tree, as a developer runs it — the lifecycle suite’s pepsi-setup run would otherwise switch to packaged system accounts the container does not have — and the three setuid/setgid suites, which need root, are then run again as root; their “not running as root” skip in the first pass is the one skip accepted. CI_ALLOWED_SKIPS (an extended regex) names any further skip accepted on purpose.

3-docs

make docs-all: every manual language including its PDF, which is the build that fails on an undeclared Unicode character.

4-upgrade

The schema-upgrade test, tests/check-upgrade.sh, against the previous release tag (a no-op until one exists): it builds that release’s pepsi-setup and pepsi-status, installs its schema and loads its frozen fixture (tests/upgrade/), upgrades with this tree taking a backup, and checks that every row survives, the backup restores, this tree’s dispatcher drains the old queue, and the old release refuses the upgraded schema with exit status 78. The script runs outside CI too, given a previous release’s binaries and SQL directory.

The Rust toolchain in the image is pinned (RUST_TOOLCHAIN in the Containerfile) rather than tracking stable, so a new clippy lint cannot turn an unrelated commit red; bump it deliberately. The jobs build into target-ci/ so they do not invalidate a developer’s target/. Run a job locally from the top of the tree, with the vendor/taler-rust submodule initialised, e.g. contrib/ci/ci.sh 2-check.

25.1.2. The lifecycle suite and its database

pepsi-test-stages spawns the real pepsi-dispatch binary against a real PostgreSQL and asserts on pepsi.workqueue row states, so it needs a reachable server. It has two database modes, chosen by PEPSI_TEST_DATABASE:

unset — the default, and the development workflow

Each test CREATEs and DROPs its own throwaway database, so the suite runs in parallel. This is what a bare cargo test -p pepsi-test-stages does; the connecting role must be able to CREATE DATABASE / DROP DATABASE.

set (e.g. pepsicheck)

Every test shares that one pre-existing database — never created, never dropped, its content preserved — resetting only pepsi.workqueue on setup. The suite then must run with --test-threads=1.

tests/check-lifecycle.sh is the second mode automated: it probes for a pepsicheck database (creating it if absent), resets the schema with pepsi-setup run --reset — after a plain run first, so a freshly created database has the _v schema the drop step deletes from — and runs the suite serially against it, leaving the database in place. If neither the probe nor the creation works it skips: exit 0, which is not a pass and not a failure.

Four environment variables tune all of this:

PEPSI_TEST_DATABASE

Selects shared mode and names the database. Unset means throwaway mode.

PEPSI_TEST_PGHOST

The socket directory for the local peer-authenticated connection. Default /var/run/postgresql.

PEPSI_TEST_PGUSER

The login/peer role to connect as. Default $USER.

PEPSI_TEST_DISPATCH_LOG

Set to 1 to see the spawned dispatcher’s own logs, which the suite silences by default.

(tests/check-lifecycle.sh additionally reads PEPSI_CHECK_DB to rename the shared database it manages, default pepsicheck.)

A single test or module is cargo test -p <crate> <name::filter>, e.g. cargo test -p pepsi-common srs::.

25.1.2.1. The privilege-boundary tests

Six test files assert the database grants of The privilege split against a real server — service_role_grants, secure_link_grants, admin_grants, whitelist_grants, setup_task and the grant cases in keystore and keydisc_lifecycle. Each runs the production provision_roles and then SET ROLEs into each account to ask PostgreSQL what it may do.

They skip (with a SKIPPED line on stderr, visible with --nocapture) when any audited role is a superuser, because a superuser bypasses every grant and nothing about the boundary can be learnt. On a development machine that is the common case: the account running the tests is pepsi, which is also the pipeline’s own role, and it is usually the cluster’s superuser too. To exercise them, run the suite with pepsi temporarily not a superuser but able to switch into the other roles, from a second superuser session, and restore it afterwards:

$ psql -U some_superuser -c "ALTER ROLE pepsi NOSUPERUSER CREATEDB CREATEROLE" \
    -c 'GRANT "pepsi-ingress", "pepsi-httpd", "pepsi-telemetry", "pepsi-whitelist",
              "pepsi-crypto", "pepsi-keydisc", "pepsi-config"
        TO pepsi WITH INHERIT FALSE, SET TRUE'
$ cargo test -p pepsi-test-stages --test service_role_grants -- --nocapture
$ psql -U some_superuser \
    -c 'REVOKE "pepsi-ingress", "pepsi-httpd", "pepsi-telemetry", "pepsi-whitelist",
               "pepsi-crypto", "pepsi-keydisc", "pepsi-config" FROM pepsi' \
    -c "ALTER ROLE pepsi SUPERUSER"

INHERIT FALSE matters: it lets pepsi become each role without acquiring its privileges, so the checks on pepsi itself stay meaningful.

The administrative API’s end-to-end tests (admin_api) include one more role-dependent case: the successful CONFIG_DB write connects with options[role]=pepsi-config, so it skips (again with a SKIPPED line) when the test role cannot SET ROLE "pepsi-config". The same grant above makes it run. None of the expiry tests there — session idle and absolute expiry, token expires_at and disabled, the pepsi-httpd prune job — waits for a window to elapse: each one backdates the timestamp the database compares against now(), which the enforcing query cannot tell apart from time passing.

25.2. The documentation builds are checks too

make docs runs all of them for the source language, and is the one documentation target in the general quality gate. It is an alias for make docs-en; make docs-de and make docs-fr build the same artifacts from the translated catalogs and make docs-all builds every language. Four separate builds sit under it, and passing one says little about the others:

make docs-html-en

Sphinx HTML. Catches broken cross-references, unknown roles and malformed tables.

make man

The section-1/5 man pages through docs/rst2man.py. Catches title over/underlines that do not match the title length — a docutils error the HTML build tolerates. Source language only: the roff pages are rendered straight from docs/*.rst and never go through Sphinx, so they have no translation catalog.

make info

The GNU info manual, through Sphinx’s texinfo builder and makeinfo. Source language only, for the same reason install-info registers one pepsi.info in one info dir index.

make docs-pdf-en

LaTeX. The one most easily forgotten, and the one that fails hardest. pdflatex has no glyph for an arbitrary Unicode character and does not degrade when it meets one: it aborts with “Unicode character … not set up for use with LaTeX” and produces no PDF at all. HTML and man render the same character without complaint, so a new arrow, box-drawing character or Greek letter in the prose passes both and breaks only this build.

Every non-ASCII character used in the manual must be declared in latex_elements["preamble"] in docs/manual/conf.py. The box-drawing set used by the ASCII-art pipeline diagrams is mapped to plain ASCII on purpose: those diagrams sit inside code-block:: text, so the replacement is expanded in a verbatim environment where math mode is unavailable.

The second failure mode is a table that does not fit on a page. Sphinx renders a list-table into LaTeX’s longtable — which splits across pages and repeats the header — only past a row-count threshold; below it the table becomes a tabular, which cannot break at all. The heuristic counts rows, but what overflows a page is height, and those diverge badly for a wide table whose cells wrap: the 113-row RFC summary table splits happily while the 17-row, seven-column comparison table in Introduction does not, and runs 397 pt off the bottom of the page.

The fix is to say so explicitly rather than rely on the heuristic:

.. list-table::
   :header-rows: 1
   :widths: 20 14 14 14 14 14 14
   :class: longtable

Unlike a missing character this does not fail the build — it is an Overfull \vbox warning and a PDF with content off the page — so grep the log for Overfull \vbox and treat anything beyond a point or two as a real defect. Sub-point overshoots are ordinary typographic rounding.

When reading its output, ignore the Hyper reference … undefined warnings from the first passes — latexmk resolves forward references on later ones. The final pass, in docs/manual/_build/pdf/en/latex/pepsi.log, is the one that must be clean.

The translations are checked the same way, and the check is not redundant: a character a translator introduces — a guillemet, a German quotation mark, a non-breaking space — fails the same pdflatex build in the same way, and no English-only run would ever meet it. make docs-all is therefore a release gate rather than a per-commit one; each language’s PDF lands in docs/manual/_build/pdf/<lang>/latex/pepsi.pdf, and a German or French build additionally needs that language’s babel support installed (Debian: texlive-lang-german, texlive-lang-french).

Keeping the translations current is make update-po, which re-extracts the message templates and merges them into docs/manual/locale/<lang>/LC_MESSAGES so that changed paragraphs come back marked fuzzy and new ones untranslated. Only the .po files are under version control; Sphinx compiles them at build time.

25.3. Topology

The suite drives three real mail accounts across three hosts, described in a small INI file, tests/test-accounts.ini. That file is not part of the source tree: create it from the template with cp tests/test-accounts.ini.sample tests/test-accounts.ini and replace its placeholder accounts with hosts you control:

Role

Account (ssh target = e-mail)

Function

ADMIN

root@<mx-host>

Runs Pepsi (ingress, dispatch, httpd), PostgreSQL and systemd; also the smarthost the served domain relays through.

ALICE / BOB

alice@<served> / bob@<served>

Local accounts on the served domain. Their MTA relays through ADMIN and is trusted by it (MYNETWORKS), so their mail is local origin (the outbound path).

CAROL

test@<other>

An account on an independent host. Her mail reaches ADMIN’s public MX on port 25 as a foreign MTA — the inbound path.

DAVE

dave@<mx-host> (optional)

A local unix user on the ADMIN host that receives mail by direct local delivery into its Maildir (the local-delivery path; no onward relay). Unlike the others it is never ssh’d into directly — its mailbox is read as root over the ADMIN connection. Only 07-maildir-test.sh needs it; 01-deploy.sh provisions the dave/dave-broken users and the pepsi-maildir group when it is set.

Every host is reached over passwordless ssh; the harness installs nothing. ALICE/BOB use an mbox (/var/mail/<user>); CAROL and DAVE use a Maildir. The mailbox helpers are format-agnostic and locate a message by a run-unique token across both layouts.

Note

The other two hosts have admission policies of their own, and they bind long before Pepsi does. They are invisible in the correctness suite, which sends a handful of messages, and decisive in the benchmarks, which send hundreds of thousands. Two in particular:

  • CAROL’s MX is the receiving end of every outbound test. When it is itself a Pepsi deployment, [pepsi-ingress] CONN_RATE_PER_SECOND and CONN_RATE_BURST there default to 1 and 1 — one new connection a second from a given source range — so a relay opening a connection per message is answered 421 4.7.0 Too many connections from your address and defers. Nothing is lost; nothing is measured either. Raise them on that host before believing any outbound figure.

  • Once they are raised, what remains is that host’s own speed. If the CAROL host is much smaller than the MTA under test, every outbound measurement is that host’s figure, not Pepsi’s; the Performance chapter says how to read such figures.

  • ALICE/BOB’s MTA is the smarthost the relay tail ends at, and Postfix’s message_size_limit defaults to ten million bytes. That is why the per-stage sweep’s larger sizes cannot traverse the smarthost path at all, which the sweep reports as the next hop’s own 552 rather than guessing.

The same file carries one non-mail account, in a [demobank] section:

[demobank]
URL  = https://bank.demo.taler.net
USER = pepsi-ci
PASS = pepsi-pass

This is a bank account at the public GNU Taler demo bank, whose currency (KUDOS) is play money. 12-wallet-topup-test.sh needs it because a manual wallet top-up only tells the user where to wire money — nothing credits the wallet until that bank transfer actually happens. The harness therefore plays the user’s bank: it logs in and makes the transfer the e-mailed instructions ask for, which is the only way the test can assert that the withdrawal completes rather than merely that instructions were sent. Register an account at https://bank.demo.taler.net/ (a new one starts with KUDOS:100 and may go KUDOS:500 into debt — roughly 120 runs of that case); DEMOBANK_USER, DEMOBANK_PASS and DEMOBANK_URL in the environment override the file. No other script uses it, and with the section absent only that one credit assertion skips.

25.4. The scripts

tests/lib.sh

Shared library sourced by every numbered script. It parses the accounts INI, provides the ssh/dependency/database/service probes used by the common pre-flight, the result counters and coloured ok/no/skip output, the ADMIN-host SQL and reset helpers, the mailbox send/receive helpers, the Taler wallet helpers, the RFC 3464 DSN assertion, and — for the extended scripts — send_smtp (a raw-SMTP sender that can set the ESMTP and DSN parameters the mail command cannot) and hop_caps (an EHLO capability probe). It is not executable on its own.

01-deploy.sh

Provisions a complete Pepsi deployment on the ADMIN host from a clean checkout: service users, PostgreSQL role/database, the configuration file (including the full stage pipeline), pepsi-setup (TLS, schema, DKIM keys, DNS records), the demo payment merchant, and the systemd units. It only verifies required tools and bails with a precise list if any are missing — it never installs OS dependencies.

02-pipeline-test.sh

The core happy-path suite (T1–T6): the pay-to-send anti-spam stage (unpaid → bounce, paid → delivered, paying centrally from the runner’s Taler wallet), the language filter (blacklisted French → bounce), the outbound signing/relay path (DKIM + ARC verified at the recipient), whitelist learning with SRS envelope rewriting, and a second local user.

03-dsn-test.sh

RFC 3461 DSN extensions and bounce variants, using send_smtp: NOTIFY=NEVER suppresses the DSN (D1); ENVID/ORCPT are echoed back (D2); RET=HDRS returns headers only (D3); NOTIFY=SUCCESS with ORIGINATE_SUCCESS_DSN yields a positive DSN (D4); an NXDOMAIN recipient routes to the default bounce (D5); an unreachable relay produces a DELAY DSN (D6); and a null-sender bounce to an SRS0=… address is reverse-decoded by ingress back to the original sender (D7).

04-settings-test.sh

The per-address override layer: a pepsi-settings row changes one recipient’s outcome but not another’s (S1); an account edits its own overrides by e-mail and is acknowledged (S2); an edit to a non-EDITABLE_STAGES section is rejected and stores nothing (S3); and one DATA with divergent per-recipient overrides is split into separate workqueue rows (S4).

05-mime-test.sh

8BITMIME / SMTPUTF8 handling: an 8-bit body and a UTF-8 Subject survive the relay intact when the delivery hop advertises the extension, are downgraded (quoted-printable/base64, RFC 2047) when it does not, and a non-ASCII recipient address that the hop cannot represent is permanently failed. The script probes the hop’s EHLO first and takes the matching branch.

06-cli-test.sh

The operator CLIs, run directly over ssh: pepsi-config (get/dump/pathsub), pepsi-whitelist (add/list/remove and the dkim/signature flags), pepsi-settings (set/get/list/unset/remove), pepsi-tlsrpt (report --dry-run / prune), and the pepsi-httpd /metrics endpoint.

07-maildir-test.sh

Direct (local Maildir) delivery via the pepsi-stage-relay-to-maildir stage, which drives the setuid-root pepsi-helper-maildir-writer. A message to the local DAVE user lands in its Maildir with the prepended Return-Path/Delivered-To/Received headers (L1); a single message to one local and one remote recipient is split, delivering dave locally while alice is relayed onward (L2); and a message to a deliberately undeliverable account (dave-broken, no usable home) gives up past MAX_LIFETIME and returns a failure DSN to the sender (L3). Because [stage-local] sits after the inbound filters (… → anti-spam → local → srs), the sender is whitelisted first to clear the pay-to-send gate. The whole script skips cleanly when DAVE is unset, the deployment has no [stage-local], or the target users are missing.

11-lmtp-test.sh

LMTP local delivery (pepsi-stage-relay-to-lmtp) against a real Dovecot on the ADMIN host, so that the recipient’s Sieve script runs as the MDA files the mail: a message whose Subject carries the Sieve trigger is fileinto :create-d into a Pepsi folder (S1, which proves Sieve ran, not merely that LMTP delivered), and a message to an unknown user is rejected per-recipient by Dovecot and routed onward to NEXT_STAGE — the stage never builds a DSN itself (S2). The script deploys and configures Dovecot itself, idempotently, and wires a throwaway [stage-lmtp] section into pepsi.conf; it drives the stage’s worker directly rather than putting LMTP into the live pipeline, so it disturbs no other script. It SKIPs cleanly when Dovecot cannot be installed or the stage binary is absent.

12-wallet-topup-test.sh

The by-e-mail wallet self-service half of pepsi-stage-auto-pay (the [stage-wallet] section): a local account tops up its own per-user GNU Taler wallet by sending pepsi-wallet@<served-domain> a Subject: command. All three methods are exercised end-to-end against the demo exchange (KUDOS), with the funded local taler-wallet-cli as the genuine peer: withdraw (manual bank-wire instructions, the wire then actually made from the [demobank] account), pull (a taler://pay-pull/ URI the harness settles) and push (the harness pushes; the user accepts the taler://pay-push/ URI). Each case asserts the control message is consumed and the wallet balance grows.

13-auto-pay-test.sh

The demand-paying half: pepsi-stage-auto-pay settles a returning pay-to-send demand from the sender’s own wallet. The demand is produced genuinely — a local user sends to a paywalled local recipient (DAVE), the outbound is internet-relayed (stamping a real Pepsi-Origin), loops back into this host’s inbound MX where pepsi-stage-anti-spam holds it and e-mails the sender a real demand. That demand is replayed through the auto-pay worker. The two cases use separate senders (hence separate per-user wallets/orders): BOB’s wallet is left empty, so auto-pay forwards the demand and charges nothing (A1); ALICE’s wallet is pre-funded ahead of time (so the slow withdrawal is out of the paywall’s short 30s payment window), so auto-pay settles the demand within it — the order flips to paid, the wallet is debited, and the held message is released and delivered to DAVE (A2).

14-multi-recipient-test.sh

Recipient fan-out, the failure mode where a message addressed to several people reaches only some of them. Both cases assert positive delivery to every intended recipient. T1 sends one envelope to two remote recipients and requires both to arrive — a relay that takes rcpt_to.first() and then terminates the row passes nothing else on. T2 covers the ~/.forward fan-out: one envelope addressed to a local user who forwards (DAVE) and a second plain recipient, where the forwarded copy and the passthrough must each arrive exactly once — a stage that reduces the database row server-side without mutating the in-memory message leaves a fused successor working from the stale recipient list. T2 needs the optional DAVE account and a deployed ~/.forward stage.

15-debian-deploy-test.sh

The packaged artefact and the wizard — the real operator install path, where 01-deploy.sh builds from a checkout and hand-writes the configuration. It builds the .deb locally, installs it on the ADMIN host (so the postinst creates the accounts, roles and dpkg-statoverride bits), then runs pepsi-setup --wizard --answers for a range of scenarios, each followed by pepsi-setup run and real probe mail through the pipeline the wizard generated. It rewrites the deployed host, so it does nothing but print its plan unless PEPSI_DEPLOY_CONFIRM=yes is set.

16-whitelist-import-test.sh

The setuid pepsi-whitelist and its mailbox importer, against the installed deployment — everything about that program that only exists once it is installed. It checks the installation itself (mode 4755 owned pepsi-whitelist, the scan helper carrying no privilege bits, the table-scoped database role, the hoster list); the namespace rule, by creating two throwaway accounts and confirming one may use <login>/… and nothing else, that list does not disclose the operator’s shared whitelist, and that -c is refused; the import itself, which must yield exactly the recipients of the mail that user sent and ignore received mail; the wildcard decision, where a busy domain collapses to one *@domain row while gmail.com with just as many correspondents does not; the privilege drop, by substituting a helper that reports its own credentials and asserting the caller’s uid appears in all three id slots; --user seeding another account’s own namespace; and finally that pepsi-stage-check-whitelist consults the per-user name its {localpart} placeholder expands to. Individual checks skip when the deployment lacks the feature.

17-crypto-test.sh

End-to-end cryptography: the parts of pepsi-stage-encrypt, pepsi-stage-decrypt and pepsi-keys that only exist once the binaries are installed setuid and the key store holds a key-encryption key. It asserts the installation and the custody boundary (4750 pepsi-crypto:pepsi; the pepsi role cannot read crypto_identity.private_wrapped), identity generation for both protocols, and interoperability rather than self-consistency: Pepsi’s ciphertext is read by a stock gpg and gpg’s ciphertext is read by Pepsi. It also covers the fail-open inbound verdicts, the forged X-Pepsi-Crypto header, and ENCRYPT = required bouncing rather than ever sending cleartext.

18-wkd-test.sh

The mirror image of 17: how other people find our keys. pepsi-httpd serves the Web Key Directory for every accepted domain, out of the key store, honouring each identity’s published flag — the policy file, the direct and advanced paths, that the bytes returned import into a stock gpg, that gpg --locate-external-key finds them over real DNS and TLS, and that an unpublished identity and a domain we do not serve are both 404.

19-secure-link-test.sh

The far end of [stage-encrypt] ON_NO_KEY = secure-link: the cleartext is taken off the wire, the recipient is mailed a link and the sender a PIN. It covers the routing, the PIN going to the sender, the portal’s disclosure behaviour, the lockout after MAX_ATTEMPTS, the successful read, the read receipt, and that the database alone holds ciphertext and no PIN.

20-admin-api-test.sh

The administrative API, over the UNIX-socket ADMIN = yes listener 01-deploy.sh configures: that all three mechanisms authenticate (SO_PEERCRED, a bearer token, a password session), that a token is confined to its scopes and a missing scope is a 403, that every read-only endpoint answers JSON, and that GET /api/v1/config discloses no secret.

21-online-setup-test.sh

That the two front ends of one setup interview — the terminal pepsi-setup --wizard and the browser one under /api/v1/setup — do not drift: same questions, identifiers, kinds, order, defaults and conditions; answers round-tripping unchanged; the same values refused by both; and the same answers rendering to the same configuration. The last case rewrites the deployed configuration (restoring it afterwards) and is gated on PEPSI_DEPLOY_CONFIRM=yes; the read-only ones always run.

22-list-api-test.sh

The mailing-list compatibility gate, and the one script here that is not ours. It runs mailmanclient 3.3.5’s own suite (344 doctest statements in using.rst plus 36 test functions) and Postorius 1.3.13’s mailman_api_tests (30 files, 248 functions) against the deployed server, with only the fixture that starts GNU Mailman Core replaced — those two replacements are in contrib/compat/, versioned beside the pins. Ahead of them it checks the things whose failure would otherwise read as 250 red Python tracebacks: that the API answers on the loopback LIST_API = yes listener and refuses without the credential, that the same paths are a byte-identical 404 on the public listener, that the test-only reset hook works, and a smoke test of the resource tree including the 400-versus-404 asymmetry on /config and the decimal-versus-hex member_id split between /3.0/ and /3.1/. A suite that cannot be installed is reported as an explained skip and still fails the run: a silently skipped gate is the other way this test can lie.

24-list-post-test.sh … 27-list-archive-test.sh

The four halves of the mailing-list subsystem over real mail: a post reaching every member with its own one-click unsubscribe URI; -join, -leave and -confirm driving the subscription workflows; a bounce being attributed to a member, scored, and turned into a warning and then a removal; and the archive ingesting, threading, searching and purging what was posted.

Script 24 also carries the digest phase: posts accumulate, an issue goes out in each format to the members who asked for that format, the volume and number are what the rule says, and a member who switched digests off that morning still gets the final issue. It is a phase of that script rather than a new one, because the deployment suite is already long.

28-list-web-test.sh

The public mailing-list pages over HTTP: that every route is a byte-identical 404 on a listener without LISTS = yes; that an advertised list is on the index and an unadvertised one is not and is still reachable by URL; that the subscribe form mints exactly one pending token, subscribes nobody by itself, re-uses a live token rather than minting a second, and answers identically for a known and an unknown address; that a confirmation GET completes nothing and only the POST does; that a post is findable at its HyperKitty URL and by search with its body escaped rather than rendered; that a private archive is a 404 and not a 403; and that the RFC 8058 one-click target removes the member with a single POST even when unsubscription_policy is moderate, is idempotent, and removes nobody on a tampered token.

Its last phase covers the member tier and the owner console: registering an account and being sent exactly one verification token (a second registration re-using it rather than minting a second), verifying and setting a password, being made an owner from the command line, signing in, and then the test that matters most — an owner changes a setting in the browser and the REST API reads the same value back, which is what proves the two surfaces write one row. It also asserts that a stranger asking for the console gets the same answer as for a list that does not exist, and that a settings post without a CSRF token is refused.

Beside the numbered scripts, tests/whitelist-suid.sh covers the same setuid contract without SSH or a deployment: run as root on any host with a local PostgreSQL, it creates a throwaway database, installs the binary setuid into a suid-honouring directory and runs the same namespace/privilege-drop assertions. make check runs it (and the other *-suid.sh scripts) best-effort — they SKIP rather than fail when they cannot get root, a database or a suid-capable filesystem.

29-list-import-test.sh

The migration tooling. A real protocol-0 config.pck imports and the conversions that are not copies are asserted through the database rather than through the converter’s return value — the archive_policy collapse, the DMARC precedence rule, the ninety days of grace period, the subject_prefix’s trailing space, the DisableMime bit producing a plaintext digest member and a bounce-disabled member arriving disabled. The import exits non-zero, because the fixture deliberately carries a ban nobody can compile, and a second run changes nothing and says so.

Its middle phase is the one worth the most: pepsi-archive export followed by import into a second list, comparing message count, every Message-ID and the thread count. That tests the two halves against each other rather than against our expectations, and a re-import adds nothing — resumption keys on content, not on a timestamp.

The last phase is the cutover mailing: the per-domain histogram, a rated run, and a second run that skips what the first invited.

What it does not do is stand up a real upstream Mailman 3 to import over REST; that needs a Mailman installation on the deployment host. The gap is named in the script rather than left to be discovered.

30-secretary-test.sh

Confirm-to-send (pepsi-stage-secretary). The deployment 01-deploy.sh builds has no secretary in its inbound chain, so the script splices one in after check-whitelist for its own duration and restores the configuration on exit. It asserts that an unknown sender is held and sent a null-sender challenge (X1); that an automatic reply to the challenge releases nothing and whitelists nobody (X2); that a genuine reply delivers the held message and whitelists the sender (X3), whose next message then passes without a second challenge (X4); that a bounced challenge sends the held mail down the timeout path at once (X5); and that with no reply the held mail is bounced after HOLD_TIME, a late reply to the expired cookie doing nothing (X6).

25.5. Running

# One-time: deploy this checkout onto the MX host (over SSH):
tests/01-deploy.sh tests/test-accounts.ini

# Then, repeatedly, from the test runner:
tests/02-pipeline-test.sh tests/test-accounts.ini
tests/03-dsn-test.sh      tests/test-accounts.ini
tests/04-settings-test.sh tests/test-accounts.ini
tests/05-mime-test.sh     tests/test-accounts.ini
tests/06-cli-test.sh      tests/test-accounts.ini
tests/07-maildir-test.sh  tests/test-accounts.ini   # needs the optional DAVE account
tests/11-lmtp-test.sh     tests/test-accounts.ini   # deploys Dovecot on the MX host
tests/12-wallet-topup-test.sh tests/test-accounts.ini   # needs taler-wallet-cli on the MX host
tests/13-auto-pay-test.sh     tests/test-accounts.ini   # needs DAVE + taler-wallet-cli
tests/14-multi-recipient-test.sh tests/test-accounts.ini   # T2 needs DAVE + a ~/.forward stage
tests/16-whitelist-import-test.sh tests/test-accounts.ini   # needs root on the MX host
tests/17-crypto-test.sh   tests/test-accounts.ini   # needs the setuid crypto stages + gpg
tests/18-wkd-test.sh      tests/test-accounts.ini   # needs pepsi-httpd + gpg
tests/19-secure-link-test.sh  tests/test-accounts.ini
tests/20-admin-api-test.sh    tests/test-accounts.ini   # needs the ADMIN = yes listener
tests/22-list-api-test.sh     tests/test-accounts.ini   # needs the LIST_API = yes listener + PyPI
tests/24-list-post-test.sh    tests/test-accounts.ini
tests/25-list-command-test.sh tests/test-accounts.ini
tests/26-list-bounce-test.sh  tests/test-accounts.ini   # needs the DAVE account for the owner notices
tests/27-list-archive-test.sh tests/test-accounts.ini
tests/28-list-web-test.sh     tests/test-accounts.ini   # needs the LISTS = yes listener
tests/29-list-import-test.sh  tests/test-accounts.ini   # the importers and the cutover
tests/30-secretary-test.sh    tests/test-accounts.ini   # splices in a secretary stage

# These two rewrite the deployed host, so they only plan unless confirmed:
PEPSI_DEPLOY_CONFIRM=yes tests/15-debian-deploy-test.sh tests/test-accounts.ini
PEPSI_DEPLOY_CONFIRM=yes tests/21-online-setup-test.sh  tests/test-accounts.ini

make integrationtests ACCOUNTS=tests/test-accounts.ini runs the whole set: the target’s INTEGRATION_SH is tests/[0-9][0-9]-*.sh minus the tests/[0-9][0-9]-*-bench.sh scripts (BENCH_SH, which have their own make benchmarks target — see Benchmark Suite), so a script is picked up by existing rather than by being listed anywhere — including any this chapter does not mention. It stops at the first failing script, and it depends on interop-thunderbird, so the S/MIME interoperability gate runs first. ACCOUNTS defaults to tests/test-accounts.ini.

Each script is self-managing: it runs a pre-flight (ssh reachability with the real error on failure, required remote tools, a database probe, and the three services’ state with the journal tail of any that is down), ensures a short PAYMENT_DEADLINE and brisk POLL_INTERVAL for fast turnaround, and resets the whitelist/queue/settings it depends on. A missing assumption fails the pre-flight with an actionable message; a message that does not arrive triggers a diagnostics dump (the live workqueue plus the dispatcher journal). Tunables such as PAYMENT_DEADLINE_SECS, DELIVER_TIMEOUT and WITHDRAW_AMOUNT are environment overrides. Each script exits 0 only when no case failed (skips are allowed).

25.6. Skips and gated cases

A case skips (rather than fails) when a precondition it cannot satisfy is absent, always with a one-line reason. The common ones:

  • No usable Taler wallet. The paid leg of 02 skips if the local taler-wallet-cli is missing or incompatible with the demo exchange. All payments are made centrally from the runner’s wallet (in production each sender would pay for their own message).

  • Deployment lacks a feature. 03/04 skip the cases that need pipeline features the deployment does not configure (ORIGINATE_SUCCESS_DSN, a [stage-internet] BOUNCE_STAGE, the [stage-delaytest] blackhole relay, the [stage-edit-settings] stage) — re-run 01-deploy.sh to enable them. 06 skips the pepsi-tlsrpt case unless a [pepsi-tlsrpt] section is configured. 07 skips entirely when the optional DAVE account is unset, the deployment has no [stage-local], or the dave/dave-broken target users are missing on the ADMIN host. 14’s T2 skips when DAVE is unset or the deployment has no ~/.forward stage.

  • Hop capability. 05 takes its pass-through, downgrade or unrepresentable-address branch according to whether the delivery hop advertises 8BITMIME / SMTPUTF8; the cases that require the hop to lack an extension skip when it advertises it. That side is covered regardless by pepsi-test-stages’ rfc2231_downgrade_lifecycle, which relays through the real smarthost stage to a mock server advertising neither extension.

  • Wallet / auto-pay not deployed. 12/13 skip entirely when the [stage-wallet]/[stage-pay] sections, taler-wallet-cli or the pepsi-wallets account are absent (re-run 01-deploy.sh). 12’s withdraw credit leg skips without a [demobank] account (see above); its pull/push cases skip without a usable local wallet. 13 skips when DAVE is unset, the genuine demand is not produced (paywalled hop or merchant down), or — for A2 — the per-user wallet cannot be funded.

25.7. Fuzzing the parsers

Three surfaces parse bytes an attacker chose — the MIME walk (pepsi_crypto::classify, which runs on every inbound message before any authentication), our own hand-written S/MIME CMS opener, and the OpenPGP reader including its compressed-packet path — plus the inbound authentication parser and the mailing-list subsystem’s parsers of bounces, posts, mboxes and imported Mailman configurations. None of them may panic, abort or run away, and that claim is kept two ways.

make fuzz runs one libFuzzer target at a time:

make fuzz                                  # FUZZ_TARGET=parse, 300 s
make fuzz FUZZ_TARGET=cms FUZZ_TIME=3600
make fuzz FUZZ_TARGET=pgp FUZZ_TIME=0      # until interrupted

The eight targets under fuzz/fuzz_targets/ are parse (the inbound authentication parser), classify (the MIME walk), cms (the CMS opener), pgp (the OpenPGP reader), bounceparse (RFC 3464 extraction and the heuristic bounce detectors), mimebuild (the parse-modify-rebuild path of the mailing-list posting stage, asserting the round trip as well as the absence of panics), pickle21 (the Mailman 2.1 config.pck reader, asserting that nothing is ever constructed) and mboxsplit (the mbox splitter the archive importer uses). FUZZ_TIME bounds a run in seconds (0 means run until interrupted) and FUZZ_FLAGS passes anything else through to libFuzzer. Seed corpora live in fuzz/corpus/<target>/, which is not tracked: PEPSI_CRYPTO_WRITE_FUZZ_CORPUS=1 cargo test -p pepsi-crypto --test sweep writes the seeds, so run it once in a fresh checkout before fuzzing; findings land in fuzz/artifacts/<target>/.

For a many-core campaign, fuzz/campaign.sh starts several targets at once under AddressSanitizer, each in libFuzzer’s fork mode with its own share of the cores, and keeps one corpus directory per target that successive runs build on:

cd fuzz && for t in parse classify cms pgp; do
    cargo +nightly fuzz build -O --debug-assertions $t; done
fuzz/campaign.sh p1 3600 "parse:24 classify:80 cms:90 pgp:54"

The first such campaign — eight hours on 256 cores, about four billion executions — found no crash, leak or hang.

fuzz/ is a separate Cargo workspace, excluded from the root one, because libfuzzer-sys needs the nightly toolchain and a sanitizer runtime — so the stable quality gates never build it, and make fuzz is a deliberate, expensive, opt-in target. Without cargo-fuzz or a nightly toolchain it SKIPs with the install hint and exits 0, so running it on a stable-only box is harmless.

The always-on half is stable and needs no nightly: pepsi-crypto/tests/sweep.rs is a deterministic mutation sweep over the same three crypto surfaces, input for input, and pepsi-common’s parse_verify_never_panics covers the fourth. Both run under plain cargo test — that is, under make check — so a regression that a fuzzing session would find is usually caught by the ordinary gate first.

25.8. Constraints

Testing must never send real mail to third-party MX hosts: the suite uses only the operators’ own domains and reserved/.invalid names (e.g. the DELAY test relays to an unrouteable 192.0.2.0/24 TEST-NET-1 address via a recipient in blackhole.invalid). The development uid cannot bind the privileged ports 25/53, so the privileged listeners are exercised only on the deployed host.