Open problems
A proposal becomes a field when other people can work on it. These are the problems, stated so that a result — positive or negative — would be publishable. Section references are to the founding paper.
1. The per-item improbability instrument (the gating problem)
The corrected inference rule needs P(shared idiosyncrasy | independent genesis) to be measurable per character. Today it is measurable only for planted marks (improbable by construction) and reference-based teacher tests. Until it is measurable for naturally occurring habits, no natural character can be called conjunctive. Everything downstream gates on this. Attribution accuracy does not substitute: discriminability is not improbability. (§4, §7.1)
2. The stemma-reconstruction benchmark
Reconstruct the descent tree of a model population from output evidence alone, validated against lineages that are a matter of record — base-model → fine-tune families, teacher → distilled students, checkpoint successions. Honest difficulty stated in the paper: public “known phylogenies” are themselves imperfect ground truth, and multi-teacher distillation is contaminatio formalised — measuring where reconstruction breaks is part of the result, not a failure of the experiment. (§7.1, §11)
3. Character selection without hindsight
A discipline for choosing candidate characters before seeing which ones group models — pre-registration, defined search spaces, multiple-testing correction. Without it the method can manufacture lineage signal, the computational analogue of selecting “significant errors” after reading the witnesses. Raised independently by an adversarial reviewer; adopted as an open problem. (§4)
4. The independence instrument for review panels
Multi-model review panels assert seat independence from vendor identity. The correlated-errors literature says vendors are not automatically independent; no instrument yet measures error correlation between specific seats on real workloads without contaminating the measurement. One proposed design was refused by a three-lineage panel — its error-only ledger could not see the correlated true findings it existed to warrant — and the refusal record is itself instructive reading. (§5, §7.2)
5. Conservation and the sealed vintage
Which model artifacts, at what fidelity, must be archived for the tradition to remain readable — weights, corpora manifests, sampling configurations? The paper’s own concession bounds the value: sealed vintages decorrelate the inheritance channel only; attractors shared before sealing stay shared. What conservation buys, and for whom, needs working out. (§7.3)
6. The diachronic thesis, as a falsifiable claim
The model ecology is becoming a textual tradition through measured channels. What observation would refute it? Candidates the paper commits to: inheritance channels weakening rather than strengthening; convergence to shared attractors swamping lineage structure entirely; the per-item instrument (problem 1) proving unbuildable in principle. A field that cannot lose is not a field. (§6, §11.14)
7. The witness with a fixed emission space
A model that returns only typed decisions over an asker-supplied option set cannot raise an unoffered reading: the option set is the whole brief (§11.6) and nothing unoffered is ever emitted (§11.7). Its probability map can signal misfit but cannot name it. Two instruments are needed. First, a per-item test of whether a flat or renormalised map tracks genuinely incomplete option sets, against oracle-labelled omissions, a calibration and out-of-distribution measurement on a closed action set that does not depend on any vendor. Second, a comparison of how closely the witness’s probability maps sit to the graders’ mean, set against competent models known not to be trained on those graders, over a full benchmark under a frozen rule, rather than an argmax side-count on a curated slice, which cannot separate “no favourite” from “follows the mean”. Proximity alone does not show descent: any competent model tends to fall between two competent graders. The comparison measures behavioural resemblance; training ancestry still needs engineered markers or reference-based teacher tests. A negative result on either changes what a panel may claim about such a seat. (§4, §7.2, §11.6–7; type specimen and counts in Field Note 1.)
Corrected 17 Sep 2026: the second instrument previously called proximity to the graders’ mean a discriminating test of descent. It is not, without comparison against competent models not trained on the graders.
Corrections to this list are welcome — the strongest contribution to a young field is showing one of its problems is ill-posed. See provenance for how to reach the maintainer.