85.1.28. pepsi-stage-detect-language

detect the human language of a message body

Manual section:

1

85.1.28.1.1. Name

pepsi-stage-detect-language - the language-detection stage of the Pepsi pipeline.

85.1.28.1.2. Synopsis

pepsi-stage-detect-language [GLOBAL-OPTIONS] worker

85.1.28.1.3. Description

pepsi-stage-detect-language is a stage program run by pepsi-dispatch(1) as a persistent worker reading message ids on standard input. It loads that pepsi.workqueue row (refusing to act unless its status is running), reads its [stage-<stage>] section, identifies the human language(s) used in the message body, and records them under state.language before advancing the message to NEXT_STAGE. This stage never drops, bounces or pauses a message, and it preserves the rest of state.

The stage extracts the message’s inline body text with a full MIME parser (shared logic in pepsi-common): text/plain parts are taken verbatim and text/html parts are reduced to plain text, each decoded from its content-transfer-encoding and its declared character set to UTF-8. Attachments (including text attachments) are ignored, as are encrypted or otherwise opaque bodies (for example multipart/encrypted or application/pkcs7-mime), which carry no readable text. For a multipart/alternative the plain alternative is preferred over the HTML one — except when that plain alternative turns out to hold no prose at all, the shape a bulk mailer produces when it puts two bare URLs where the message should be; the HTML sibling is then used instead.

85.1.28.1.3.1. Prose filtering

The extracted text is then reduced to the part of it that is prose, and only that is classified. A statistical detector models how humans write words, and an e-mail body is full of things that are not words: URLs, addresses, base64 keys, DNS records, hexadecimal identifiers. Unfiltered, a body that is mostly machine text is classified from the machine text: a DNS-record notification or a spam body of two bare URLs gets a confident verdict for a language nobody wrote in it, even when every word a human wrote there is English.

Two filters run, in this order:

  • tokens that cannot be words are dropped — anything containing @, a slash or backslash, or an internal full stop, anything beginning http or www., anything less than half letters, anything containing a digit, anything longer than 30 characters, and anything whose capitalisation changes inside the word (CamelCase, base64);

  • lines left with fewer than 15 letters, or without a letter majority, are dropped — what filter one leaves of a DNS dump or a table is a scatter of bare field names, which is not prose in any language.

The rules are aware of scripts that do not separate words with spaces (Han, Kana, Hangul, Thai, Khmer, Lao, Burmese, Tibetan): there a whole clause arrives as one token, so the length, digit, full-stop and capitalisation rules are not applied to it, and lines are judged by how many letters they hold rather than how many words. Without that, a Chinese or Thai message would be deleted in its entirety.

85.1.28.1.3.2. Classification

The surviving prose is classified with the lingua statistical detector, built from the candidate languages named by the LANGUAGES option. The detector yields a probability for each candidate language; the most likely at most three languages whose probability is at least 5 % are recorded.

At most the first 20 000 characters of that prose are classified. The detector scores every candidate language over the whole text, so its cost is the text’s length times the number of candidate languages — and with LANGUAGES = * that is 75 of them, against a message size limit measured in megabytes. A language is decided by the first few kilobytes; classifying the rest of a long message would only occupy the worker. There is no option for this: it is a bound on the work, not a policy.

If the message has no usable inline text, none of that text is prose, or there is too little prose (fewer than 20 characters) to classify reliably, the stage skips classification: the message advances unchanged, with no language key added. Recording nothing is deliberate — a body that is only URLs has no language, and guessing one from a URL slug is how a message ends up confidently tagged with a language nobody wrote it in.

85.1.28.1.4. Output format

state.language is a string in the syntax of an HTTP Accept-Language header: a comma-separated list of code;q=quality items, ordered from most to least likely. Each code is the language’s ISO 639-1 code and each quality is its detection probability in the range 0–1 (trailing zeros trimmed, so a certain single language reads q=1). For example, a message that is mostly English with some German might be recorded as:

en;q=1, de;q=0.45

The detector is expensive to construct, so each worker process builds it once (per distinct candidate-language set) and reuses it for every message. It is built at lingua’s full accuracy; the library’s low-accuracy mode is deliberately not used. That mode short-circuits — it returns a language at probability 1 as soon as exactly one candidate owns an n-gram occurring anywhere in the text, without examining the rest — so a single address in a quoted attribution line, or a base64 key, can decide the language of kilobytes of prose. On filtered prose it is not cheaper either: the shortcut search is paid for and then discarded, and it preloads extra n-gram count models, so it buys nothing.

85.1.28.1.5. Configuration

Options live in the stage’s own [stage-<name>] section (PROGRAM = pepsi-stage-detect-language): a NEXT_STAGE and the candidate LANGUAGES set. They are documented in pepsi.conf(5).

85.1.28.1.6. State

Inputs: none from state; the stage reads the stored message body.

Outputs: when at least one language is detected, the stage merges language (the Accept-Language-style string) into state and advances — one update that moves the row to NEXT_STAGE and merges the key; otherwise state is left untouched and the row simply advances. The state layout is described in pepsi.state(7).

Transitions: always advances to NEXT_STAGE, whether or not a language was detected. There is no branch; the stage never pauses, fails, splits or finishes. A message whose content makes the classifier panic is advanced without a language, with a warning in the log: retrying it would only panic again, and nothing downstream requires a language.

85.1.28.1.7. Commands

worker

Run as a persistent pepsi-dispatch(1) worker, reading message ids on standard input.

85.1.28.1.8. Global Options

-c FILE, –config FILE

Read the configuration from FILE instead of searching the default locations.

-L LOGLEVEL, –log LOGLEVEL

Set the logging verbosity (default info).

-v, –verbose

Show log messages from all sources.

-h, –help; -V, –version

Print a usage summary / the version and exit.

85.1.28.1.9. Exit Status

0

The message was processed (advanced, with or without a recorded language).

1

An error occurred (message not found or not running, or a per-address setting that breaks the stage’s section — an invalid LANGUAGES code, say). The reason is written to the log.

A fault of the host (the database, a template or helper that cannot be used) is not reported as a failure: the message is paused and retried, as Stage errors in pepsi-dispatch(1) describes. A section that does not parse makes the worker refuse to start (status 78) instead of failing each message in turn. So is a missing NEXT_STAGE.

85.1.28.1.10. Diagnosis and tuning

pepsi-detect-language(1) is the offline counterpart of this stage: the same executable under a second name, which classifies a message read from a file or standard input and prints only what this stage would have recorded, without going near the database. Use it to reproduce a misclassification from a saved message (--explain shows the extracted text, the prose left after filtering it, and every candidate’s confidence) and to try candidate sets before configuring one (--languages).

85.1.28.1.11. Examples

Classify one message by hand (the id must name a running row):

echo 42 | pepsi-stage-detect-language -c /etc/pepsi/pepsi.conf worker

A pipeline section that tags the body language before the whitelist check:

[stage-detect-language]
PROGRAM = pepsi-stage-detect-language
NEXT_STAGE = check-whitelist
LANGUAGES = en de fr es

85.1.28.1.12. See Also

pepsi-detect-language(1), pepsi-config(1), pepsi-stage-check-whitelist(1), pepsi-dispatch(1), pepsi.conf(5), pepsi.state(7), pepsi-setup(1)

85.1.28.1.13. Bugs

Report bugs to the Pepsi issue tracker.