Agentic Stemmatics
the philology of machine-generated text

The philology of machine-generated text

Large language models emit testimony — fluent claims whose warrant lies outside the act of generation. Agentic stemmatics is the proposal that the right science for such output is the one philology built over five centuries of handling corrupt textual transmission: stemmatics, the discipline that reconstructs lost originals and hidden lines of descent from the errors and idiosyncrasies of surviving witnesses.

The mapping is not decorative. Model outputs are witnesses sampled from lineages. Training corpora are exemplars. Distillation is copying. Contamination of benchmarks and training data is contaminatio. Watermarks are deliberately planted Leitfehler — the marked readings by which a copy betrays its source. Model collapse under recursive training is the same distributional decay philology observes when hard readings are banalised through generations of copying. Each of these research areas exists today with its own literature; the field-move is the unifying frame and the inference discipline that comes with it.

The corrected inference rule

The naive borrowing — shared errors imply shared ancestry — does not transfer to models, and the field’s founding paper spends much of its length demonstrating why: on the measured evidence, shared model errors are dominated by convergent attractors, not inheritance. What survives is a corrected rule:

Model ancestry rides shared idiosyncrasies that are improbable under independent genesis — checked per item, never granted to a channel wholesale.

The honest scope, stated first

Because calibration is this field’s signature discipline, its boundary belongs on its front page: the corrected rule is presently operative for engineered markers and reference-based teacher tests only. Per-item improbability for naturally occurring model habits is not yet measurable, so natural-character stemmatics is a programme, not a result. The paper claims a framing, a corrected inference rule, and a specified experimental agenda — not demonstrated efficacy. What would make the programme exact, and what would falsify it, are stated as open problems.

Why it matters now

Machine text re-enters the archive: model output is scraped into training corpora, distilled into students, and diffused through human writing. If the model ecology is becoming a genuine textual tradition — a falsifiable conjecture, not an assumption — then questions philology answers about manuscripts become askable, and eventually answerable, about models: what did this model learn from? which models are copies of which? what has been lost from the distribution, and can it be conserved? The stakes range from scientific (provenance and evaluation) to practical (multi-model review panels whose independence is asserted but unmeasured) to archival (model vintages as sealed witnesses for a future that will want to read this period’s record).

Start with the founding paper, work through what’s open, or see how existing literatures map onto the frame.