Adjacent literatures
The field’s constituent research mostly exists — scattered across ML, security and digital humanities venues, without a shared frame. This map states what each area is, stemmatically, with an entry point into its literature. It is an orientation, not a survey; each row deserves — and in most cases has — a literature of its own.
| research area | stemmatically, this is | entry point |
|---|---|---|
| Watermarking & antidistillation fingerprinting | Deliberately planted Leitfehler — synthetic marked readings by which a copy betrays its exemplar. The demonstrated inheritance of planted marks is the cleanest existing proof that the descent channel is real. | Kirchenbauer et al. 2023; the antidistillation sampling literature |
| Distillation & teacher detection | Copy identification: was this witness produced by reading through that exemplar? Reference-based tests are one of only two presently clean character classes. | Rawat et al. 2026 (arXiv:2607.09692) |
| Model collapse / recursive-training degradation | Transmission decay: distributional tails erode under repeated copying, as hard readings are banalised across scribal generations. The diachronic thesis’s best-measured channel. | Shumailov et al., Nature 631 (2024); Gerstgrasser et al. 2024 |
| Benchmark & training-data contamination | Contaminatio — readings crossing between lines of transmission, the corruption philology calls hardest to detect. Ambient contamination is the ecology’s default condition, not its exception. | the data-contamination literature; Kobak et al. on LLM traces in scientific prose |
| Correlated errors across models | The refutation of the naive borrowing: shared errors are dominated by convergent attractors, not ancestry — which is precisely why the corrected rule rides idiosyncrasies instead. Load-bearing for multi-model panel design. | Kim et al., ICML 2025 (arXiv:2506.07962) |
| Authorship attribution & stylometry for LLMs | The usus scribendi instrument: per-lineage habit profiles. Caution the field insists on: attribution accuracy measures discriminability, which is not the improbability condition ancestry inference needs. | Juzek & Ward 2025 (“Why does ChatGPT ‘delve’ so much?”) |
| Human-mediated diffusion of model idiom | Transmission through the human population back into corpora — the channel that makes the archive’s future strata machine-inflected regardless of lab policy. | Spennemann 2026; corpus studies of post-2022 lexical shift |
| Classical stemmatics & digital philology | The donor discipline — five centuries of method for exactly this problem shape, including its own internal critique (Bédier’s objection to over-neat trees), which transfers too. | Maas, Textkritik; Trovato 2014 |
Two boundaries the map keeps visible. Presence in a corpus is not descent — showing machine text re-enters the archive is weaker than showing character states are heritably transmitted, and only planted marks currently demonstrate the stronger claim. And shared method is not shared ancestry — models converge on the same behaviours because they optimise similar objectives on overlapping data; the corrected rule exists because the naive one fails exactly here.