.. This file is part of PEPSI. Copyright (C) 2026 GNUnet e.V. PEPSI is free software; you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation; either version 3, or (at your option) any later version. ======== Archives ======== **This archive is a reimplementation of HyperKitty**, the archiver of GNU Mailman 3, whose data model, threading rules, Message-ID hash and URL scheme it follows. HyperKitty is copyright the Free Software Foundation and its contributors, and upstream's documentation describes the same concepts; the compatibility is deliberate and load-bearing, not incidental. A site migrating from Mailman keeps its archive links, and importing its archive is close to a table copy. What differs is where the archive lives: in PostgreSQL, alongside the queue and the roster, rather than in a search index plus a document store. That is what makes the thing this chapter opens with possible. The archive is one half of a subsystem: the other half is the lists themselves. :doc:`mailing-lists` is the chapter for lists, members, moderation and digests, and :ref:`list-architecture` there is the map of the whole thing — where the archive is written from (the ``to-archive`` handler of :doc:`programs/pepsi-stage-list-post`) and what else reads it. The operator's tool is :doc:`programs/pepsi-archive`; the browser surface is :doc:`web-ui`; importing an existing archive is part of a migration, so see :doc:`installation`; the standards are collected in :doc:`rfc-index` (RFC 5064 ``Archived-At`` in particular); and :doc:`features` summarises what is and is not implemented. Anything the REST API says about archiving is :doc:`mailman-api`. GNU Mailman itself is acknowledged at greater length in :doc:`mailing-lists`, which names the project, the Free Software Foundation and the GNU General Public License the imported material stays under. Purging a message ================= **Deleting a message from a public archive is one command.** .. code-block:: console # pepsi-archive purge --list announce@lists.example.org \ 'CAFxyz...@mail.example.org' purged VHBHHV4YTF64QIJBGZ6XWA57YD6VRIZA from announce@lists.example.org; tombstoned for 90 days so a resend or a re-run import cannot bring it back The message, its attachments and its votes are gone; the thread it was in is renumbered; the counters are repaired; and an ``event_log`` row records who purged what and when. **It does not delete the Message-ID.** A *tombstone* keeps it for ``[pepsi-list] TOMBSTONE_RETENTION`` days (default 90), and while that tombstone stands the archive refuses to store that message again. There are exactly two ways a purged message comes back — somebody re-delivers it, or an importer is run a second time — and both are accidents rather than decisions. ``purge --forget`` skips the tombstone when the decision is deliberate. That is a departure from "delete means gone", and it is stated here rather than in a footnote: for ``TOMBSTONE_RETENTION`` days after a purge, the archive still holds that message's identifier, its purge date and the name of whoever purged it. What is archived, and what is not ================================= A list's ``archive_policy`` decides who may read the archive: ``public``, ``private`` (members only) or ``never`` (nothing is stored at all). The first two are enforced **at read time**, which is what lets an owner flip a list from private to public without a re-ingest — and what makes ``purge`` the only thing that ever removes content. A sender may opt one message out with an ``X-No-Archive:`` header (whatever its value) or ``X-Archive: no``. **The archive keeps the real sender address.** Obscuring it is a rendering decision: an anonymous reader sees an obscured form, a signed-in member sees the address. Upstream publishes addresses as posted; we store them in full and decide on the way out, because the importer, the exporter and mbox round tripping all need the real value — and because a setting that destroys data on the way *in* cannot be changed back. Searching ========= Two mechanisms, deliberately both, because neither subsumes the other. **Ranked word search** stems: a search for ``upgrade`` finds ``upgraded``. It uses the list's own language dictionary, and where PostgreSQL has none it says ``simple``, which still matches — it just does not stem. A hit in the subject outranks one in the body, and the sender's name is searchable. **Substring and fuzzy search** does not stem, and that is what it is for: a hostname, a config key, a function name, a line from a traceback, a misspelled name. On a technical list that is a large share of what people actually search for, and a stemming dictionary throws exactly that material away. **The search box decides which one to use.** A word or a quoted phrase is a word search; a fragment carrying punctuation (``lists.example.org``, ``--no-certbot``, ``self.assertEqual``) or an explicit ``*wildcard*`` is a substring search; and the result says which one answered. The trigram tier ---------------- Substring search costs index space, so how much of a list is indexed for it is a choice — ``off``, ``short`` or ``full``: .. list-table:: :header-rows: 1 :widths: 15 85 * - Tier - What is indexed for substring search * - ``off`` - Nothing. The column is ``NULL``, and PostgreSQL's GIN index does not index NULLs — so a list that opts out costs **exactly nothing**, which looks like an omission and is the whole mechanism. * - ``short`` - The subject, the thread's subject, the sender's name and address. The default: enough for "that message from Alice about releases". * - ``full`` - The above plus the message body. The expensive one, and the only one that finds a fragment of a traceback. ``[pepsi-list] SEARCH_TRIGRAM`` is a **site ceiling**, not a default; a list chooses up to it with ``pepsi-list list set-ext search_trigram ``. The effective tier is the lower of the two. Changing it is a **reindex, not a migration**: .. code-block:: console # pepsi-archive reindex --list announce@lists.example.org announce.lists.example.org: reindexed 12043 message(s) at tier full with the english dictionary Lowering the site ceiling therefore *shrinks* the index at the next reindex rather than merely forbidding new rows. **A list at ``short`` says so.** A body-fragment search against it answers "this list does not index message bodies for substring search" rather than nothing at all: a search that quietly looks at less than the reader thinks it does is worse than one that admits it. A site-wide search reads every partition (the archive is partitioned by list), so it is the slower, explicit choice; a search scoped to one list prunes to one partition and uses that list's own dictionary. URLs a migration keeps ====================== Every archived message has a *Message-ID hash*: base32 of the SHA-1 of its ``Message-ID``. It is upstream's, unchanged, because every ``Archived-At:`` header any HyperKitty deployment ever emitted points at a URL derived from it, as does every link in every mail anyone kept: =================================================== ========================= URL What it names =================================================== ========================= ``/archives/list//`` the list's archive ``/archives/list//message//`` one message ``/archives/list//thread//`` a thread, by its starter =================================================== ========================= A thread's identifier is its **starting message's** hash, which is why threading has to match upstream's rather than merely be sensible: a thread grouped differently is a thread with a different URL. Threading ========= The parent of a message is its ``In-Reply-To``, falling back to the **last** entry of ``References`` — the message actually replied to, not the thread's root. If that message is not archived on this list, the reply **starts a thread of its own**. That is upstream's rule and it is kept, so an imported archive threads the way it did before; the repair is a separate command rather than a different rule: .. code-block:: console # pepsi-archive rebuild-threads --list announce@lists.example.org announce.lists.example.org: reattached 37 orphaned repl(ies), renumbered 1284 thread(s) A reply loop — two messages that claim each other as parent, which real mail contains — drops one edge rather than hanging. Counters, and the one way to break them ======================================= The index pages read counter **columns** (messages per thread, when a thread was last active) rather than counting a million rows on every render. They are maintained by the writer, which means one thing an operator has to know: **If you write to the archive tables outside ``pepsi-archive``, the counters lie.** They are repairable, which is why they are columns rather than a cache: .. code-block:: console # pepsi-archive recount --list announce@lists.example.org Retention and export ==================== ``pepsi-archive expire --list --before `` deletes everything older than a date, in batches — the predicate is a date and the partitioning is by list, so it cannot prune, and one statement over a year of a busy list would hold a transaction open for minutes. The same deletion runs on a schedule. ``pepsi-list tasks --once``, which the **pepsi-list-tasks.timer** unit runs daily, expires each list's messages older than its retention: a list's own ``archive_retention_days`` (set with ``pepsi-list list set-ext``) wins, ``0`` included; a list without one uses ``[pepsi-list] ARCHIVE_RETENTION``, whose own default is ``0``; and ``0`` keeps everything. So an archive keeps everything unless somebody decides otherwise, threads and counters are repaired as after a manual ``expire``, and a list can keep everything on a site that expires. A stored value that is not a day count is skipped with a warning rather than guessed at, since a wrong guess deletes mail. ``pepsi-archive export --list `` writes an ``mboxrd`` mbox to standard output, which is the format every other archiver reads. It is ``mboxrd`` rather than ``mboxo`` on purpose: a body line that already began ``>From `` gets another ``>``, which makes the export **invertible** — and an invertible export is the difference between a backup and an approximation. Attachments =========== An attachment is user-supplied bytes served from our own origin, which is the most dangerous thing in this subsystem. Four rules, applied by the archive rather than by whatever renders the page: * always ``Content-Disposition: attachment``, never inline; * a stored ``text/html`` part is **never** served as ``text/html``; * the stored content type is advisory — the response uses a safe type from a short allow-list (plain text, PNG, JPEG, GIF, WebP, PDF) and ``application/octet-stream`` for everything else, so a type nobody thought about is boring rather than dangerous; * the response carries a ``Content-Security-Policy`` of ``default-src 'none'; sandbox`` (plus ``base-uri``, ``form-action`` and ``frame-ancestors`` all ``'none'``) and ``X-Content-Type-Options: nosniff``. Browsing it =========== The archive is served on a listener flagged ``LISTS = yes`` (see :ref:`three-listener-flags`), at **HyperKitty's URL shapes**: .. list-table:: :header-rows: 1 :widths: 55 45 * - URL - Page * - ``/archives/list//`` - Overview: every month with a count, and the most recently active threads * - ``/archives/list////`` - One month's threads * - ``/archives/list//thread//`` - A whole thread, indented by stored depth * - ``/archives/list//message//`` - One message * - ``/archives/list//message//attachment//`` - One attachment * - ``/archives/list//search?q=…`` - Search this list **The shapes are preserved on purpose.** ```` is the same Message-ID-Hash, so every ``Archived-At:`` header a migrated deployment ever emitted still resolves — including the ones already sitting in subscribers' mailboxes and quoted in other people's archives. Keeping them costs a route table and is worth more than any improvement to them. Who may read what ----------------- An anonymous visitor sees a list's archive only when its ``archive_policy`` is ``public``. Anything else is a **404**, not a 403: saying "forbidden" would confirm that the list exists, and an unadvertised list is reachable by URL precisely so that the difference matters. Every browse page is scoped to one list, so that decision is taken once before any query runs. **Search is the exception**: it spans lists and *ranks* them, and a rank computed over rows the viewer cannot see leaks their existence through the ordering of the ones they can — so there the visibility rule is part of the ``WHERE`` rather than a filter applied afterwards. Attachments ----------- An attachment is always served as a **download**, never as the type the message claimed for it, and under a policy that lets it do nothing if a browser decides to render it anyway. This is one helper with four rules in it, and the web pages link to it rather than serving anything themselves; a test asserts that no code on the public surface sets a content type from stored data. Two documented differences from upstream ---------------------------------------- Both are visible to a reader and neither is accidental: * **Addresses are obscured for anonymous viewers** (``alice@...``). The archive keeps the real address; this is a rendering decision, and it is not a security control — see :doc:`web-ui`. * **``archive_rendering_mode = markdown`` is stored and not rendered.** The attribute exists because the attribute set is the compatibility contract, and both values render as text. Rendering user-supplied markdown into our own origin is what these pages' Content-Security-Policy exists to make impossible. Reader actions that need an account ----------------------------------- Two things on an archive page belong to a person rather than to the page, and both need a member account (see :ref:`two-account-systems`): ``POST /archives/list//message//vote`` A vote of ``1``, ``-1`` or ``0`` to withdraw. Keyed on ``(list, message, member)``, so it is idempotent by construction — pressing the button twice is one vote — and the answer is a redirect back to the message, so a reload cannot re-vote either. The votes are stored per **message** and the counters live on the **thread**, which is HyperKitty's arrangement: a thread list shows a score, and computing it per page load would mean a join over every message's votes. The counters are therefore recomputed from the vote rows in the same statement as the vote, rather than incremented — a reader who changes an up vote to a down vote moves the count by two, and an increment that assumed one would drift. ``POST /archives/list//thread//favourite`` A toggle, because the control is one button and a button has no state to send. Both carry a CSRF token, and an anonymous reader asking for either is sent to the sign-in form: the archive itself stays readable without an account, and only these writes need one. A signed-in member also sees the **whole sender address** rather than the obscured form. That is not a weakening of the rule above — obscuring was never a security control, and anybody on the list already has the address at the top of their own copy. What it stops is the bulk harvest that makes a public archive a spam source, and a harvester does not register an account and verify an address to read one page at a time. The consequence to know about is a caching one: a page whose content depends on who is reading it is served ``no-store``, so an archive read by signed-in members is not shareable by a front cache. The archive has no thread tagging and no categories. ``pepsi-archive import`` reads mbox files with the repairs in the reader and rebuilds the threads afterwards; see :doc:`programs/pepsi-archive` and :doc:`installation`. ``pepsi-archive``'s command line is the whole of the archive's operator interface.