85.1.28. pepsi-stage-detect-language¶
detect the human language of a message body
- Manual section:
1
85.1.28.1.1. Name¶
pepsi-stage-detect-language - the language-detection stage of the Pepsi pipeline.
85.1.28.1.2. Synopsis¶
pepsi-stage-detect-language [GLOBAL-OPTIONS] worker
85.1.28.1.3. Description¶
pepsi-stage-detect-language is a stage program run by pepsi-dispatch(1)
as a persistent worker reading message ids on standard input. It loads that pepsi.workqueue row
(refusing to act unless its status is running), reads its
[stage-<stage>] section, identifies the human language(s) used in the message
body, and records them under state.language before advancing the message to
NEXT_STAGE. This stage never drops, bounces or pauses a message, and it
preserves the rest of state.
The stage extracts the message’s inline body text with a full MIME parser
(shared logic in pepsi-common): text/plain parts are taken verbatim and
text/html parts are reduced to plain text, each decoded from its
content-transfer-encoding and its declared character set to UTF-8. Attachments
(including text attachments) are ignored, as are encrypted or otherwise opaque
bodies (for example multipart/encrypted or application/pkcs7-mime), which
carry no readable text. For a multipart/alternative the plain alternative is
preferred over the HTML one — except when that plain alternative turns out to
hold no prose at all, the shape a bulk mailer produces when it puts two bare URLs
where the message should be; the HTML sibling is then used instead.
85.1.28.1.3.1. Prose filtering¶
The extracted text is then reduced to the part of it that is prose, and only that is classified. A statistical detector models how humans write words, and an e-mail body is full of things that are not words: URLs, addresses, base64 keys, DNS records, hexadecimal identifiers. Unfiltered, a body that is mostly machine text is classified from the machine text: a DNS-record notification or a spam body of two bare URLs gets a confident verdict for a language nobody wrote in it, even when every word a human wrote there is English.
Two filters run, in this order:
tokens that cannot be words are dropped — anything containing
@, a slash or backslash, or an internal full stop, anything beginninghttporwww., anything less than half letters, anything containing a digit, anything longer than 30 characters, and anything whose capitalisation changes inside the word (CamelCase, base64);lines left with fewer than 15 letters, or without a letter majority, are dropped — what filter one leaves of a DNS dump or a table is a scatter of bare field names, which is not prose in any language.
The rules are aware of scripts that do not separate words with spaces (Han, Kana, Hangul, Thai, Khmer, Lao, Burmese, Tibetan): there a whole clause arrives as one token, so the length, digit, full-stop and capitalisation rules are not applied to it, and lines are judged by how many letters they hold rather than how many words. Without that, a Chinese or Thai message would be deleted in its entirety.
85.1.28.1.3.2. Classification¶
The surviving prose is classified with the lingua statistical detector, built from the candidate languages named by the LANGUAGES option. The detector yields a probability for each candidate language; the most likely at most three languages whose probability is at least 5 % are recorded.
At most the first 20 000 characters of that prose are classified. The
detector scores every candidate language over the whole text, so its cost is the
text’s length times the number of candidate languages — and with LANGUAGES =
* that is 75 of them, against a message size limit measured in megabytes. A
language is decided by the first few kilobytes; classifying the rest of a long
message would only occupy the worker. There is no option for this: it is a bound
on the work, not a policy.
If the message has no usable inline text, none of that text is prose, or there is
too little prose (fewer than 20 characters) to classify reliably, the stage
skips classification: the message advances unchanged, with no language
key added. Recording nothing is deliberate — a body that is only URLs has no
language, and guessing one from a URL slug is how a message ends up confidently
tagged with a language nobody wrote it in.
85.1.28.1.4. Output format¶
state.language is a string in the syntax of an HTTP Accept-Language
header: a comma-separated list of code;q=quality items, ordered from most to
least likely. Each code is the language’s ISO 639-1 code and each quality
is its detection probability in the range 0–1 (trailing zeros trimmed, so
a certain single language reads q=1). For example, a message that is mostly
English with some German might be recorded as:
en;q=1, de;q=0.45
The detector is expensive to construct, so each worker process builds it once
(per distinct candidate-language set) and reuses it for every message. It is built
at lingua’s full accuracy; the library’s low-accuracy mode is deliberately not
used. That mode short-circuits — it returns a language at probability 1 as
soon as exactly one candidate owns an n-gram occurring anywhere in the text,
without examining the rest — so a single address in a quoted attribution line,
or a base64 key, can decide the language of kilobytes of prose. On filtered prose
it is not cheaper either: the shortcut search is paid for and then discarded, and
it preloads extra n-gram count models, so it buys nothing.
85.1.28.1.5. Configuration¶
Options live in the stage’s own [stage-<name>] section
(PROGRAM = pepsi-stage-detect-language): a NEXT_STAGE and the candidate
LANGUAGES set. They are documented in pepsi.conf(5).
85.1.28.1.6. State¶
Inputs: none from state; the stage reads the stored message body.
Outputs: when at least one language is detected, the stage merges
language (the Accept-Language-style string) into state and advances —
one update that moves the row to NEXT_STAGE and merges the key; otherwise
state is left untouched and the row simply advances. The state layout is
described in pepsi.state(7).
Transitions: always advances to NEXT_STAGE, whether or not a language was detected. There is no branch; the stage never pauses, fails, splits or finishes. A message whose content makes the classifier panic is advanced without a language, with a warning in the log: retrying it would only panic again, and nothing downstream requires a language.
85.1.28.1.7. Commands¶
- worker
Run as a persistent pepsi-dispatch(1) worker, reading message ids on standard input.
85.1.28.1.8. Global Options¶
- -c FILE, –config FILE
Read the configuration from FILE instead of searching the default locations.
- -L LOGLEVEL, –log LOGLEVEL
Set the logging verbosity (default
info).- -v, –verbose
Show log messages from all sources.
- -h, –help; -V, –version
Print a usage summary / the version and exit.
85.1.28.1.9. Exit Status¶
- 0
The message was processed (advanced, with or without a recorded language).
- 1
An error occurred (message not found or not
running, or a per-address setting that breaks the stage’s section — an invalid LANGUAGES code, say). The reason is written to the log.A fault of the host (the database, a template or helper that cannot be used) is not reported as a failure: the message is paused and retried, as Stage errors in pepsi-dispatch(1) describes. A section that does not parse makes the worker refuse to start (status 78) instead of failing each message in turn. So is a missing NEXT_STAGE.
85.1.28.1.10. Diagnosis and tuning¶
pepsi-detect-language(1) is the offline counterpart of this stage: the same
executable under a second name, which classifies a message read from a file or
standard input and prints only what this stage would have recorded, without going
near the database. Use it to reproduce a misclassification from a saved message
(--explain shows the extracted text, the prose left after filtering it, and
every candidate’s confidence) and to try candidate sets before configuring one
(--languages).
85.1.28.1.11. Examples¶
Classify one message by hand (the id must name a running row):
echo 42 | pepsi-stage-detect-language -c /etc/pepsi/pepsi.conf worker
A pipeline section that tags the body language before the whitelist check:
[stage-detect-language]
PROGRAM = pepsi-stage-detect-language
NEXT_STAGE = check-whitelist
LANGUAGES = en de fr es
85.1.28.1.12. See Also¶
pepsi-detect-language(1), pepsi-config(1), pepsi-stage-check-whitelist(1), pepsi-dispatch(1), pepsi.conf(5), pepsi.state(7), pepsi-setup(1)
85.1.28.1.13. Bugs¶
Report bugs to the Pepsi issue tracker.