85.1.53. pepsi-telemetry

anonymous feature-telemetry collector

Manual section:

1

85.1.53.1.1. Name

pepsi-telemetry - central, anonymous, opt-in feature-telemetry collector.

85.1.53.1.2. Synopsis

pepsi-telemetry [GLOBAL-OPTIONS] serve

85.1.53.1.3. Description

pepsi-telemetry is the central collector for Pepsi’s anonymous, opt-in feature telemetry. It is a small HTTP server, deliberately separate from pepsi-httpd(1) and shipped in its own Debian package, intended to run on a single host (by default telemetry.pepsi.taler.net) behind an existing nginx or Apache front server that terminates TLS and rate-limits.

Pepsi installations that opt in ([pepsi] SHARE_TELEMETRY = YES; the option defaults to no, so an installation that never set it submits nothing and is never even assigned a system_id) submit two kinds of report, identified only by a random 256-bit system_id (no personally identifying information). The collector stores them in the pepsi.telemetry table, counting usage per reported Pepsi version, and exposes an aggregate report that never reveals an individual system_id.

The collector connects to the shared pepsi PostgreSQL schema as the pepsi-telemetry role over the local socket. On a host that only serves telemetry the rest of the schema simply stays empty. Install and upgrade it with pepsi-setup -c /etc/pepsi-telemetry/pepsi-telemetry.conf schema (see pepsi-setup(1)); like every Pepsi program, the collector refuses to start (exit status 78) against a schema from another release.

85.1.53.1.4. Endpoints

All three endpoints live under /telemetry/. The two submission endpoints are open and unauthenticated — the only deployment identity is the anonymous system_id, and that identity is self-asserted: the collector neither issues it nor checks anything about it beyond its 64-hexadecimal-character shape, so it is an identifier, never a credential. Abuse is bounded by a request-body cap (MAX_BODY), the per-submission limits described below, and the fronting web server’s rate limiting — note that the shipped Apache site uses mod_ratelimit, which throttles response bandwidth and does not limit the request rate at all; only the nginx site has limit_req.

POST /telemetry/features

Record a deployment’s enabled-feature snapshot. Body:

{"system_id": "<64 hex>", "version": "0.1.0",
 "features": {"arc": true, "dane": "strict", ...}}

Each feature value must be a JSON scalar: true for a plain on/off feature, or a string/number for a feature carrying a mode. The snapshot is authoritative for (system_id, version): features it lists become enabled (their scalar kept as a detail value), features it omits are marked disabled (their accumulated usage is retained). Replies 204 on success, 400 on a malformed body / bad system_id / non-scalar value, 413 when the body exceeds MAX_BODY, 408 when the body does not finish arriving within 30 seconds.

Three further per-submission bounds apply to both submission endpoints, and exist because this is the one component exposed to the internet: at most 512 names in one body, each name 1 to 64 characters of [A-Za-z0-9._-], and a version of at most 32 characters (it is part of the row key, so an unbounded one is an unbounded row supply). Every name Pepsi reports is a compile-time constant of that shape, so the charset can be this narrow.

POST /telemetry/usage

Accumulate feature-usage counts. Body:

{"system_id": "<64 hex>", "version": "0.1.0",
 "usage": {"arc": 42, "srs": 7, ...}}

Each value is a non-negative integer no greater than 2^40: the number of times the feature was exercised since the previous submission (a delta). Counts are summed server-side and attributed to the reported version, and the running per-deployment total is clamped at 10^12 — the endpoint is unauthenticated and nothing prunes a row, so a count near the 64-bit ceiling would otherwise overflow the aggregate below and wedge it permanently. The optional period_start/period_end fields are accepted and ignored. Replies as for /telemetry/features.

GET /telemetry/report

Return the public aggregate report as JSON. One entry per feature with the cross-version rollup — deployments and uses (total exercised count) — plus a per-version by_version breakdown. deployments counts the deployments whose stored snapshot has the feature enabled, and “enabled” means “exercised at least once since pepsi-telemetry-client(1) last started” — not “configured”. Two consequences: a feature that is switched on but never triggered (dane where no destination publishes TLSA, tlsrpt on a quiet day, a bounce stage that never fires) is never reported enabled at all; and because the daemon’s set is per-process, restarting it flips that deployment’s whole feature set to disabled until each feature is exercised again, so a report generated in that window undercounts. No system_id ever appears, so the report is safe to serve openly. Example:

{"generated": "2026-06-27T12:00:00Z",
 "features": [
   {"name": "arc", "deployments": 298, "uses": 91234,
    "by_version": [{"version": "0.2.0", "deployments": 210, "uses": 80012}]}
 ]}

The contrib/update-feature-stability.sh script turns this report into the manual’s feature-stability table.

The figures are unverified, and cannot be otherwise as the protocol stands. Because system_id is self-asserted (above), deployments counts distinct identifiers that were submitted, not distinct installations: anybody can mint a hundred random identifiers and post a feature snapshot and a usage count under each, which is by itself enough to carry a feature over the stable thresholds of the generated table. Read the report, and the stability tiers derived from it, as adoption reported in good faith rather than as evidence. Closing the gap needs the collector to issue the identifier at a first-submission handshake and return a secret that later submissions present (compared in constant time), or an equivalent signed-submission requirement.

A missing or empty version is recorded as the empty string rather than rejected, so a client that reports none still contributes (bucketed as an unknown version).

85.1.53.1.5. Configuration

The [pepsi-telemetry] section configures the single listener and the request limits; the database connection comes from the shared [pepsi-postgres] section.

Always pass -c /etc/pepsi-telemetry/pepsi-telemetry.conf. That file is where the shipped configuration lives, but it is not the default: the collector’s configuration-source component is pepsi, so with no -c the usual search finds /etc/pepsi/pepsi.conf or /etc/pepsi.conf — the mail pipeline’s configuration, which on a collector-only host does not exist at all. pepsi-telemetry.service passes the option, which is why the shipped deployment works; a run by hand has to.

SERVE

systemd (what the shipped configuration sets, socket-activated through pepsi-telemetry.socket), unix or tcp — the same listener vocabulary the other servers use. In unix mode set UNIXPATH (/run/pepsi-collector/pepsi-telemetry.sock, matching what the shipped reverse-proxy sites proxy to) and optionally UNIXPATH_GROUP (the front web server’s group, e.g. www-data) and UNIXPATH_MODE (default 0660). tcp mode (BIND_TO/PORT, default port 8080) serves plaintext and is intended for local testing only. There is no TLS in pepsi-telemetry itself; the fronting web server terminates it.

MAX_BODY

Maximum accepted request-body size in bytes (default 65536).

MAX_CONNECTIONS

Cap on simultaneously-served connections (default 128).

DB_POOL_SIZE

Database connection-pool size (default 2).

85.1.53.1.6. Global Options

serve is the only subcommand. These global options precede it (a trailing flag is rejected).

-c FILE, –config FILE

Read the configuration from FILE instead of searching the default locations ($XDG_CONFIG_HOME/pepsi.conf, $HOME/.config/pepsi.conf, /etc/pepsi/pepsi.conf, /etc/pepsi.conf) — none of which is the collector’s own file, so always pass it; see above.

-L LOGLEVEL, –log LOGLEVEL

Set the logging verbosity. LOGLEVEL is one of error, warn, info, debug or trace (default: info).

-v, –verbose

Show log messages from all sources, including third-party libraries.

-h, –help

Print a usage summary and exit.

-V, –version

Print the version and exit.

85.1.53.1.7. Files

/etc/pepsi-telemetry/pepsi-telemetry.conf

The collector’s configuration, shipped by the Debian package. It is not found automatically — see above — so every invocation needs -c /etc/pepsi-telemetry/pepsi-telemetry.conf.

/run/pepsi-collector/pepsi-telemetry.sock

UNIX socket the collector listens on: bound by pepsi-telemetry.socket under the shipped SERVE = systemd, or created by the daemon itself if you set SERVE = unix with this UNIXPATH. The directory comes from the service’s RuntimeDirectory=pepsi-collector (with RuntimeDirectoryPreserve=yes, so stopping the service does not delete a socket still bound inside it), or from systemd when it binds the socket.

The directory belongs to this service alone, and both halves of that are deliberate. It is not /run/pepsi, the mail pipeline’s, and it is not /run/pepsi-telemetry, the telemetry client’s (whose own socket is /run/pepsi-telemetry/socket, owned by pepsi). systemd chowns a runtime directory and its contents to the declaring unit’s User= on every start, so two units sharing one directory under different accounts rewrite each other’s sockets: were the collector’s socket in the client’s directory, the client starting would take the socket’s www-data group away and the front server would answer 502 for every request while the collector stayed running and logged nothing, having never seen a connection. Distinct file names in a shared directory do not help; a directory of one’s own does.

/etc/nginx/sites-available/telemetry.pepsi.taler.net, /etc/apache2/sites-available/telemetry.pepsi.taler.net.conf

Reverse-proxy site files shipped (disabled) by the Debian package; enable the one matching your front server.

85.1.53.1.8. See also

pepsi-httpd(1), pepsi-setup(1), pepsi.conf(5).