CounterProof Research · Paper
CounterProof Research, a practice of Clavestra Capital Limited. Preprint, not peer-reviewed. Published version of record: doi:10.5281/zenodo.22030516 (CC BY 4.0).

Agentic Stemmatics: Collation, Provenance, and Conservation for an Emerging Machine Textual Tradition

A field report and a research programme

L. J. Soons · Director and co-founder, CounterProof Research, a practice of Clavestra Capital Limited (Malta, C 113987) · ORCID: 0009-0000-3088-6373 · Correspondence: admin@counterproof.io

Concept DOI 10.5281/zenodo.22030516 resolves to the latest deposit. This text is v1.37, a revision of v1.35 that corrects the record of its review; v1.35 was deposited on 25 Sep 2026 as version DOI 10.5281/zenodo.22959463 (deposited file SHA-256 84c621ce603dd843ba3e10e4a997ef8316720752febb0a922aa78fc60447a437); its full revision history is in Appendix A. External claims carry [verify]; internal anecdotes are labelled illustrations, not evidence.

Competing interest (full statement after the Abstract): the author is the director and co-founder of CounterProof Research, a practice of Clavestra Capital Limited, which sells adversarial review services using the method described here and would benefit from its adoption; every internal observation in this paper comes from that practice.

Abstract

A language model can make a claim without supplying good grounds for believing it. Fluency alone cannot tell us whether those grounds exist. This paper proposes treating model outputs as testimony and draws on stemmatics, the branch of philology that studies relationships among surviving textual witnesses.

The comparison requires an important correction. Shared errors are common in the 2023–24 model population, but they do not establish that one model inherited those mistakes from another: models can reach the same plausible mistake without copying, and the available measurements do not show how many did. Our argument concerns how this situation is changing. Model outputs return to training corpora through four measurable channels: direct distillation, unlabelled inclusion in collected text, deliberate synthetic-data production, and human writing influenced by models. Through these channels, we argue, the ecology is growing a genuine descent tradition, in which later models inherit material from earlier ones; its extent and growth are unmeasured here.

On the rule this paper proposes, tracing that inheritance requires shared features, singly or jointly, unlikely to have arisen independently. We therefore investigate distinctive habits and markers in model output, with attribution and distillation-detection research supplying candidate instruments. Engineered markers provide demonstrated cases; the value of naturally occurring habits for reconstructing ancestry remains to be established.

From this framework we develop a conditional argument about review: majority agreement is unsafe as a means of establishing truth where reviewer independence is unmeasured or failing. Research on correlated errors indicates this condition is the present norm, with greater correlation among more capable models. We propose an experiment to reconstruct model relationships from shared features that are improbable under independent origin, and to test whether this recovers known lineages better than existing trees built from aggregate output similarity. We also propose preserving model vintages and provenance records, with explicit rules for responding when review panels become less independent.

The paper reports failures of our own instruments, including the earlier draft of this argument and the recorded responses to its reviewers. It also reproduces the frozen commitments for our first controlled comparison, which was discontinued on critique, before any run occurred, and records two dated events from that instrument (§§8–9). First, a disputed claim that internal agreement alone could not settle, asserted by three commissioned investigations, was substantially settled, with edge corrections, by primary sources: a specification and two public issue records. Second, when three external reviewing lineages split over how the instrument's governing clause applied, the split was recorded rather than smoothed over, and the decision to discontinue did not depend on either reading. The contribution is a framework, a corrected inference rule, and a research programme. Its effectiveness has not been demonstrated.

Competing interests

The author is the director and co-founder of CounterProof Research, a commercial practice of Clavestra Capital Limited (Malta, C 113987) that provides and sells adversarial review services using the method described here. Every internal observation in this paper comes from that practice, including the field record in §8, the commitments in §9, and the review history in §12. The practice stands to benefit from adoption of the method. No external funding supported the work.

This relationship directly shapes the limitations discussed in §11: selection and publication bias, the unaudited internal record, and the author's role in adjudicating his own panels. We state the interest before the argument so that readers can assess both together.

This paper was drafted with an AI assistant of the Claude lineage (Anthropic), which also prepared the review briefs and the collations of the reviewers' returns; the author directed the work and is responsible for it. None of the explicitly identified reviewing seats in this paper's record is attributed to Claude. The record also includes reviews whose lineage is unconfirmed or unspecified. Anthropic has publicly alleged that several AI developers used Claude's outputs at scale to improve their own models, among them the developers of four lineages that have served as seats here: Moonshot (Kimi), DeepSeek, Alibaba (Qwen) and Zhipu (GLM) [Anthropic, "Detecting and preventing distillation attacks," 23 February 2026]; [Anthropic, "Detecting and countering misuse of AI: September 2026," 10 September 2026, "Illicit distillation"]. These are an interested party's allegations and are not examined here; §12 states what they could mean for this paper's review.

1. Introduction: model outputs as testimony

Consequential judgments are increasingly delegated to language models: whether a program is correct, whether a vulnerability is real, whether a finding can be closed. A fluent answer may be well supported or mistaken, and the ease with which we read it does not tell us which. Where the stakes are high, the common safeguard is to ask several models and to treat their agreement as confirmation. That safeguard rests on an assumption the measurements do not support. Models share their mistakes; when two are wrong they are often wrong in the same way, and the more capable they are, the more their errors correlate (§4). Agreement is being counted as evidence exactly where it is weakest. This paper asks what discipline can weigh testimony from witnesses that are fluent, fallible and partly dependent on one another, and finds it in a branch of philology built for that situation.

A conventional program can, in principle, be examined to understand how its instructions produce its behaviour. Reading a model's answer serves a different purpose: it tells us what the model asserts. The truth of that assertion is not reliably recoverable from its fluency or internal consistency.

The same generative mechanism produces supported and unsupported claims. It assigns probabilities to possible continuations; the probability of a token is not the probability that a claim is true. Whether internal features reliably distinguish the two remains a question for interpretability. The reader cannot assume that the answer itself supplies such a distinction [Turpin, Michael, Perez & Bowman, "Language Models Don't Always Say What They Think," NeurIPS 2023, arXiv:2305.04388; Lanham et al., "Measuring Faithfulness in Chain-of-Thought Reasoning," Anthropic, 2023].

We therefore treat model outputs as testimony. Strictly speaking, the witness is the model lineage, sampled repeatedly; each answer is one draw from it (§3). Its statements must be examined, compared, and traced to their sources. Agreement matters, but its value depends on the relationships among the witnesses who agree.

Philology has long studied this problem. Surviving manuscripts may preserve an earlier reading, repeat a common corruption, or combine material from several sources. Stemmatics investigates these relationships to reconstruct how texts were transmitted. It offers a disciplined way to reason about witnesses that are fallible and partly dependent on one another.

The comparison with models has limits. Two models may arrive at the same wrong answer without either having copied the other. Similar training data and objectives can lead both towards the same plausible mistake. Shared error alone therefore does not establish shared ancestry. On the rule this paper proposes, a feature, or a set of features assessed jointly, supports that inference only when independent appearance in both models would be unlikely; §4 states the rule as proposed, to be tested, not as an established result. Determining which features meet this condition is a central task of the programme.

An earlier version of this paper failed to respect that limit. Version 0.1 called the relationship between stemmatics and model review "structurally identical" and treated the method as directly transferable. Four separately commissioned reviewers unanimously returned it for major revision. Their two central objections were correct: the paper had performed no distinctively stemmatic operation, and it had mistaken shared model errors, which can arise without descent, for evidence of descent. The present argument incorporates those corrections. It makes a weaker claim about the present and a stronger, testable claim about the tradition emerging as model outputs enter later training corpora. It also proposes an experiment whose results can be checked against known model relationships.

The paper makes seven contributions:

  1. A mapping with stated limits. Section 3 sets out where stemmatics applies to model review and where the comparison fails.
  2. A corrected rule for inferring ancestry. Section 4 explains the problem of independently arising similarities, known as homoplasy. It examines shared idiosyncrasies as possible evidence of descent and asks which satisfy the proposed requirement that independent origins be unlikely. It distinguishes the measured evidence of correlated error from the evidence still needed for particular lineage signals.
  3. A conditional argument about aggregation. Section 5 examines majority voting when reviewer independence is unmeasured or failing. The argument rests on external research into correlated errors. The costs and failure modes of our own alternative, adjudication, receive the same scrutiny.
  4. An account of recursive transmission. Section 6 examines how model outputs return to training corpora, and what this may mean for tracing origins, conserving earlier material, and understanding security risks.
  5. A research programme. Section 7 outlines an experiment to reconstruct model ancestry, a proposed measure of reviewer independence, and an archival approach with rules for responding to declining independence.
  6. A field record of failure. Sections 8 and 12 describe failures of our instruments, including this paper's earlier draft and a dated external instance from a sibling artifact. These cases motivate the framework; they do not demonstrate its effectiveness.
  7. A record of an unrun comparison. Section 9 reproduces the frozen commitments for our first controlled comparison, preserves their restrictions on what could be claimed, and explains why the instrument was discontinued before execution.

The purpose of this programme is constructive. Machine-generated work may be reliable enough for consequential use when the surrounding process supplies adequate grounds for relying on it. Those grounds must be established through examination and verification. They cannot be inferred from the finished prose alone.

More capable models may produce a higher proportion of sound claims. That improvement does not make fluency a certificate of truth: supported and unsupported claims still arrive through the same channel (§6). Yet fallibility and shared errors do not, by themselves, justify excluding machine-generated work from every consequential task. The question is whether we can specify an assurance process, measure its performance, and expose its failures. Sections 3–5 and 7 begin that work.

This paper cannot itself establish that such a process succeeds. Its own review history is subject to the limits discussed in §§12–13. In particular, we do not claim that the method produces better code. The controlled comparison intended to test that claim was a registered experiment whose commitment values existed, but it was discontinued before any run occurred (§9).

2. Stemmatics and its conditions

Many classical texts survive only as copies of copies. The original, or archetype, has been lost. Stemmatics reconstructs archetypal readings by comparing the agreements and disagreements among the surviving witnesses.

The method is conventionally associated with Karl Lachmann. That attribution requires care. Timpanaro argues that Lachmann adapted and formalised an existing philological practice, rather than inventing it. Its predecessors included the humanist tradition of emendatio ope codicum, Bentley, and F. A. Wolf's Prolegomena. The dispute concerns the method's originality, not its existence [Timpanaro, The Genesis of Lachmann's Method, trans. Most, University of Chicago Press, 2005; originally La genesi del metodo del Lachmann, 1963].

The central inference rests on the conjunctive error: a mistake so particular that its independent appearance in two manuscripts would be unlikely. Such an agreement therefore evidences common ancestry. Two conditions govern the inference.

Agreement must be unlikely to have arisen independently

Not every shared error carries evidence of descent. Independent scribes can simplify a difficult passage, make the same easy slip, or modernise the same spelling. Philology excludes these polygenetic errors precisely because they can arise more than once.

Maas called the significant errors Leitfehler, by analogy with geology's index fossils, and distinguished separative from conjunctive errors [Maas, Textkritik, 1927; English Textual Criticism, trans. Flower, Clarendon, 1958]. The rule was never simply that a shared error proves descent. A shared feature supports descent only when its independent origin would be improbable. We express that condition as:

P(shared feature | independent genesis) is negligible.

This notation is our attempt to make the criterion operational for computation. Maas expressed it qualitatively as unwahrscheinlich, or improbable; he did not use this probability notation.

Biology distinguishes the corresponding cases as homoplasy, or independently arising similarity, and synapomorphy, or a shared derived character. Classical stemmatics already required the logical distinction, although its application depended on the editor's judgement. Modern computational stemmatology adds formal machinery, including character-state matrices and cladistic parsimony. It shares that machinery with phylogenetics and tests it against artificial traditions whose true relationships are known [Roelli, ed., Handbook of Stemmatology, De Gruyter, 2020; Roos & Heikkilä, "Evaluating methods for computer-assisted stemmatology using artificial benchmark data sets," 2009].

The method must carry its own criticisms

Stemmatics has never enjoyed an uncontested success. Bédier found that 105 of 110 published stemmata had two branches. He concluded that the method partly manufactures the structure it claims to find [Bédier, 1928].

Two responses must be distinguished. Bédier turned towards editing from the best manuscript. He retained the aim of recovering a single text while abandoning genealogy as the means of doing so. The later New Philology took another course: it made textual variation itself the object of study and relinquished the search for a single archetype [Cerquiglini, In Praise of the Variant, trans. Wing, Johns Hopkins, 1999; originally Éloge de la variante, 1989]. Earlier drafts of this paper wrongly presented these as one continuous development.

Both criticisms matter here. A method may inspire confidence in a structure that it cannot fully verify from within. Our comparison with Bédier concerns this risk. It is a structural analogy, not a claim that he studied model panels. Section 11 considers the corresponding difficulty of establishing a panel's independence from inside the panel itself.

Four further terms recur. Their definitions against standard handbooks remain [verify]:

We do not claim that manuscripts and language models are identical kinds of things. We borrow a discipline for reasoning about transmission, corruption, dependence, and independently arising agreement. Its terms are useful only where their meanings survive the comparison. Where they do not, as with strict eliminatio and emendatio, we retain whatever logical lesson remains and relinquish the name. Section 3 makes those limits explicit.

3. Where the comparison holds

Three rows of the earlier draft's mapping that reviewers showed to be false or forced have been corrected or removed. The following account replaces them. Each correspondence has a stated status, and the corrections change the argument rather than merely its wording.

The lost archetype and the question to be settled

Status: conditional. The analogue of the archetype is ground truth about an artifact, but only where the question can be settled by an appropriate oracle. Throughout this paper, decidable has this operational meaning. An execution, type-check, test, proof, or consultation of a normative text may settle the relevant question. The general property is undecidable in the computability-theoretic sense; this bounds the framework's domain.

The limits of each oracle matter. Reading a specification can establish which enumeration it contains. Reading an issue record can establish its recorded status. Neither act establishes that the document is correct about the world. It establishes what the document says.

Questions of severity or design soundness may admit several legitimate judgements. Disagreement need not mean that someone has corrupted a single recoverable truth. Even where a definite answer exists, our situation differs from the philologist's: our archetype is often extant but expensive. We can read the code or run a test. Wherever an appropriate oracle can settle the question, comparison among witnesses is subordinate to it (§11).

The manuscript witness and the model lineage

Status: corrected. A single answer is not the witness. The witness is the model lineage understood as a distribution and characterised through repeated sampling. One answer is one draw from that distribution.

The previous draft contradicted itself by treating an output as a witness while acknowledging that a model returns different verdicts on identical input. The correction has practical consequences. A blocking finding is a sampled claim to be established against the artifact; it is not itself a verdict (§5; §11.3).

Shared errors and shared ancestry

Status: replaced. The classical inference from a significant conjunctive error does not transfer wholesale to shared model errors, because shared errors are common and can arise without copying, and the available measurements do not show what share did (§4). On the corrected rule it transfers only where independent origin is improbable, for an error alone or together with other shared features, and such an error counts as one significant shared idiosyncrasy among others. The corrected rule, which this paper proposes and §7.1 would test, is an inference from significant shared idiosyncrasies, singly or jointly, with significance determined by the improbability of independent origins.

Redundant witnesses and repeated samples

Status: renamed. Classical eliminatio codicum descriptorum removes a copy because it adds no genealogical information. A second sample from a model family is different: it can reveal the lineage's variability.

We therefore discount samples according to measured correlation rather than eliminate them. The Latin term no longer fits the operation. Its warning nevertheless remains useful: correlated witnesses must not be counted as though they were independent.

Contamination between reviewers

Status: a narrow correspondence. Giving one reviewer another's findings resembles copying from a second exemplar. This is illustrated, not evidenced, in our field record (§8); review survival is not offered as warrant (§13).

A common evidence bundle is a milder relative: common inputs rather than mixed exemplars. Reviewers share inputs without necessarily borrowing one another's conclusions. Common inputs and mixed exemplars should be recorded separately.

Independent agreement and confidence

Status: conditional on measured independence. Agreement can strengthen a claim when the relevant dependence has been assessed. Different vendors do not by themselves establish independent errors (§§4–5). Nor can a panel fully establish its own independence from within (§11). Where this paper says reviewers worked separately or blind, it describes a procedure: none saw another's answer. That is not a measured independence of their errors.

The harder reading

Status: removed as a rule for choosing between present outputs. The suggestion that we should distrust the more fluent answer gives us no defined measure of difficulty. Scribal error is not uniformly directional either [verify].

The comparison returns at a different level in §6, where as a claim about iterated transmission it becomes exact: recursive training on model output erodes the tails of a distribution, just as scribal banalisation erodes difficult readings [Shumailov et al., Nature 631:755, 2024]. This concerns the development of a tradition through time; it does not tell us which of two current answers is true.

The shared intermediate ancestor

Status: a functional analogy. A classical hyparchetype is a particular lost intermediate manuscript that may be reconstructed. A shared training corpus is no such single ancestor. We extend the term to name a risk: witnesses may inherit blind spots from a common source, leaving comparison among them unable to reveal what all omit. This use goes beyond the term's strict codicological meaning (§6; §11).

Several of these lessons are already familiar to reliability engineering as common-mode failure. Their value does not depend on Latin terminology. The additional contribution of philology lies in its account of descent, historical layers, mixed transmission, and conservation, and in a worked historical example of reconstruction from outputs without access to the mechanisms that produced them. That example establishes that output-level science proved adequate for reconstruction-grade warrant without mechanism-level access in its own domain. Whether comparable warrant is possible here is part of what §7 must show (§10).

The criticism that the earlier draft performed no distinctively stemmatic operation therefore stands. Section 7 answers it by outlining experiments; it does not claim that those experiments have been performed.

A final boundary follows from the first correspondence. A generator can enlarge the range of available answers. Deciding which question matters remains a judgement, and no quantity of fluent answers settles it.

4. Similarity without descent

The strongest objection to the earlier draft struck at its central inference. A conjunctive scribal error supports descent because the coincidence would be improbable without copying. Shared model errors need not have that character. Two models trained on overlapping data towards similar objectives can arrive at the same plausible falsehood without any copying between them.

This is the problem of homoplasy. A similarity may have several origins, and its appearance alone does not tell us which one occurred.

What the error measurements establish

Kim, Garg, Peng and Garg studied more than 350 models across two leaderboards and a résumé-screening task. When two models were both wrong, they gave the same wrong answer at rates far above chance: a mean of 60% on one leaderboard and 42% on the other. More capable models showed more correlated errors ["Correlated Errors in Large Language Models," ICML 2025, PMLR 267:30038–30066, arXiv:2506.07962].

Pairs from different providers were included throughout the study. Its statistical model attributed more agreement to capability than to shared vendor on one leaderboard, and the reverse on the other. Most agreement was attributed to neither. The study did not report an analysis restricted to cross-provider pairs. These qualifications limit what we can infer about vendor diversity specifically.

A second, capability-controlled audit formed all 31,900 two-, three- and four-model subsets of 30 models on MMLU-Pro; the subsets overlap and are not independent observations. Among the 4,060 three-model subsets, majority vote beat the best individual member, identified after the fact, in 9.98%; when the best member was chosen on held-out items, the rate was 18.71% (mean over 20 random splits). After capability control, its author found that more shared error goes with lower gain, and described that association as modest and its magnitude as configuration-dependent [Kim, "Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles," arXiv:2607.20768, 2026; preprint, not peer-reviewed at the time of writing].

A third study measures a judging panel directly. Nine frontier models from seven families, labelling three natural-language-inference datasets, supplied about two independent votes' worth of information, and the best single judge matched or outperformed the full panel. Same-family pairs were only slightly more correlated than cross-family pairs, the three most correlated pairs were cross-family, and a panel of one judge per family did not recover independence [Kohli, "Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels," arXiv:2605.29800, 2026; preprint, also published by Apple Machine Learning Research]. The task is classifying text, not reviewing code, so the result bears on review panels by analogy rather than by direct measurement.

We infer that consensus selection can discard correct minority answers when errors overlap. The audit's finding that majority vote rarely beat the best individual member identified after the fact, and the panel study's finding that the best single judge matched or outperformed the full panel, are consistent with it, but an aggregate comparison does not count the correct minority answers that voting discarded. What these results establish for our genealogical argument is that shared error is common. They do not separate errors that arose without copying from errors that were transmitted. Philology's exclusion rule does not require that separation, because it asks whether independent origin of a shared feature would be improbable, not how the feature in fact arose (§2); for models the question can also be put to a set of features jointly (below).

Attribution is not yet ancestry

Models also display stable habits in vocabulary, structure, and formatting. Sun and colleagues report 97.1% five-way attribution among major systems, with robustness to paraphrase [Sun, Yin, Xu, Kolter & Liu, "Idiosyncrasies in Large Language Models," ICML 2025, arXiv:2502.12150].

That result shows that the systems can be distinguished. It does not establish the probability that a particular shared habit would arise independently. The distinction is essential: a feature useful for recognising a model need not reveal its ancestry.

The overuse of words such as "delve" supplies a caution. It appears across unrelated models after 2022 and in human scientific writing. Shared reinforcement learning from human feedback, or RLHF, is a candidate explanation for the model pattern [Juzek & Ward, "Why Does ChatGPT 'Delve' So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models," COLING 2025, arXiv:2412.11385]. The authors leave the causal mechanism open. We therefore retain it as a candidate, not an established cause.

The habit cannot presently be treated as evidence of textual descent. Moreover, a shared methodology is a different explanation from either independent attraction towards the same answer or inheritance through text. Recognising that possibility requires knowledge of how the models were built. Output comparison alone is insufficient.

Why models may converge

A language model is trained to predict the next token from its context. Under uncertainty, generation tends towards the more probable continuations. Models trained on overlapping corpora towards similar objectives can approximate similar distributions and return the same peaks. They can therefore make the same mistake without copying one another.

Deployed models are not pure seekers of the modal answer. Temperature, instruction tuning, and preference training alter their behaviour. The direction survives each of them, though that claim still requires an anchor [verify: mode preference persisting under temperature, instruction tuning, and preference tuning].

This account also explains why we treat a verdict as one draw from a witness. Repeated answers vary because generation samples from a distribution. Rerun instability is consequently relevant to the mechanism, rather than merely an oddity of a particular product (§11). The knowledge invoked here concerns the public training objective, not access to private weights.

A third origin of similarity

Shared tooling deserves its own category. Two development processes may use the same methodology, component, or vocabulary. Their products can then share habits even if no model-generated text passed between them. At a broader causal level, this and attraction towards similar distributions belong to the same family of common causes: a shared training distribution is itself a shared component of a build.

Tokenisation offers a concrete case. Models using the same vocabulary can share segmentation habits and failures on tokens that occur in that vocabulary but received little training. Land and Bartolo examine this failure class across a diverse model population ["Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models," EMNLP 2024, arXiv:2405.05417]. Our citation was checked at abstract level only [verify: full-text confirmation of the measurement's scope and model-discriminating power].

Without a category for shared tooling, the analyst may mistake such a habit for textual inheritance or dismiss it as noise. Neither response is adequate. We need an inventory of shared tokenisers, base checkpoints, and training practices. Much of this is public provenance information or can be investigated by output probes; it does not require access to weights.

The analogy with philology survives this admission. An editor need not analyse papyrus chemistry, but must know which conventions were common to the copying trade. A habit taught throughout a scriptorium cannot be treated as an improbable coincidence. Likewise, the improbability of independent model origins cannot be assessed solely from the observed resemblance (§10).

Which features can support descent

The condition must be applied to each feature, or to each set of features assessed jointly. Two classes currently stand on different evidential grounds.

Engineered markers. A statistical mark can be planted in a teacher's outputs and sought in a distilled student. Independent origin is improbable by construction. The inheritance result is due to Sander et al., "Watermarking Makes Language Models Radioactive," NeurIPS 2024, and Gu et al., ICLR 2024. With open-model access, Sander and colleagues detected inheritance when as few as 5% of the teacher-generated fine-tuning instructions were marked and the remainder came from the same teacher without marks. A separate mixture with human text as the remainder allowed detection at 10% marked instructions.

Pan et al. examine that inheritance while challenging its durability as a defence ["Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?," ACL 2025, arXiv:2502.11598]. An attacker can defeat it by recovering the marking rule, diluting the marked text, or neutralising the mark during generation. Ordinary paraphrase under a secret key does not remove it. We cite the study for the inheritance it assumes and measures, while retaining its adverse robustness result. Fragility limits durability; it does not remove the underlying improbability argument.

Naturally occurring habits. Formatting preferences, structural patterns, and characteristic glyph use require the same scrutiny as shared errors. The attribution studies examined here distinguish models, but do not measure how often each habit would arise under independent training. Stable and distinctive is not enough.

The philological term usus scribendi names characteristic writing habits used in attribution and dating. Reynolds and Wilson's Scribes and Scholars is cited as a general account of scribal transmission, not as a confirmed treatment of that precise term [4th ed., Oxford University Press, 2013; originally 1968]. Pasquali's Storia della tradizione e critica del testo (Le Monnier, 1934) and Bischoff's Latin Palaeography: Antiquity and the Middle Ages (trans. Ó Cróinín & Ganz, Cambridge University Press, 1990) are, per a philological review of this draft, more precise candidate anchors, but neither has been confirmed against the primary text [verify]. A 2019 digital-humanities project page using the expression is not offered as that primary source.

Reference-based distillation detection comes closer to the question we need to answer. It compares a student's outputs with those of a named candidate teacher, rather than merely classifying text among a fixed set of known systems [Rawat, Chen, Anand, Duan, Rotsted & Min, "Reference-Based Distillation Detection in LLMs," arXiv:2607.09692]. Section 7.1 therefore gives these methods a central place in the proposed validation experiment.

This is the paper's central theoretical claim:

In manuscripts, ancestry is carried by significant shared errors. In models, ancestry is carried by significant shared idiosyncrasies. In both cases "significant" carries its full weight: an idiosyncrasy counts only where independent genesis is improbable, as an error does. For models the condition may be met by a single feature or by a set of features assessed jointly, with their dependence modelled; it may not be granted to an entire class of outputs. That the stemmatic engine transfers only where this condition is restored and checked is the rule this paper proposes, not an established result: whether requiring the condition recovers documented lineage better than raw similarity does is the test of §7.1. Nor would the condition be sufficient: improbability under independent genesis tells against that explanation, but does not by itself identify a parent, a direction or a route, since an unsampled common source can produce the same agreement.

On the evidence considered here, shared model errors are frequent, and what share of them arose without copying is unmeasured. The rule does not exclude an error from counting: a shared error qualifies when its independent origin is improbable. Stemma provides related evidence: fingerprints built from stable, robust wrong answers that one background model does not usually reproduce help distinguish related from unrelated pairs of checkpoints (§10). That result does not establish an independent-origin probability for each selected answer. At least one prominent shared habit fails that condition on present evidence. Whether a particular failure arises through convergence or shared tooling requires information about the build processes, as well as comparison of outputs.

Two consequences follow. First, the objection to the earlier draft is itself the central, named, partially-solved problem of both source disciplines, stemmatics and phylogenetics. This is consistent with the comparison being useful, but it is not evidence that the framework is correct. To treat a refutation as confirmation would make every outcome support us (§11.13).

Second, if capability increases error correlation across vendors as well as within them, then a panel of more capable models from different vendors has more correlated errors than a comparable panel of weaker ones, whatever vendor diversity itself contributes. Kim, Garg, Peng and Garg found greater correlation among more capable models in a population that included different providers; their study did not isolate cross-provider pairs, and the association is cross-sectional, not a measured trend through time. The practical response is to measure independence in the relevant setting (§7.2).

5. What agreement can establish

The earlier draft called majority voting "the wrong aggregation function." That was too broad. Voting can improve precision, suppress false positives, and make triage affordable when true defects are rare. Under genuine independence it also yields Condorcet-style gains. Our argument concerns the conditions under which agreement fails as evidence of truth.

The conditional claim

Majority voting and consensus are unsafe as truth procedures when reviewer independence is unmeasured or failing. Correlated agreement can count the same evidence several times. A voting threshold can also discard a correct minority finding when the majority shares a blind spot.

The research discussed in §4 measures error correlation in populations that include different providers and reports greater correlation among more capable models. The further claim about losing correct minority findings is our inference from that association. The aggregate comparisons in §4 are consistent with it, but they do not count the correct minority findings that voting discarded.

This conditional does not provide a general boundary between truth-seeking decisions and operational choices. Across three revisions, we tried to construct such a boundary. The first attempt misused the literature on judgement aggregation, which is defined by what is aggregated, not by whether an aggregation tracks truth. Later attempts invoked Condorcet's Jury Theorem and applied a Popperian test recursively to a decision's premises. External review exposed defects in each version. The third still depended on separating observation from aggregation, although almost any pooled judgement can be redescribed as an observation of the pooling. We withdrew the general theory.

A narrower rule remains. If a decision depends on a factual premise established by pooling judgements whose independence was unmeasured or failing, the condition applies to that premise. Calling the overall decision "governance" does not exempt it. Ask whether the decision would change if the premise were false. Where it would not, that premise is not carrying the decision, and this condition does not apply to it. Section 9 records our own failure to respect the distinction.

The systems and their differences

The diversity assumption, that differing data, architecture, and provider produce decorrelated errors, is not established by the measurements in §4: errors agreed far above chance in a population that included different providers, and that study reports no analysis restricted to cross-provider pairs [Kim, Garg, Peng & Garg, ICML 2025]. The panel study in §4 does compare families, and finds cross-family pairs almost as correlated as same-family ones, on a text-classification task [Kohli, arXiv:2605.29800, 2026]. Yet systems using several models differ in ways that matter to this argument.

Mozilla.ai's Star Chamber queries Claude, GPT, and Gemini. It labels findings Consensus when all agree, Majority when two or more agree, and Individual observation otherwise [Wilson, "The Star Chamber: Multi-LLM Consensus for Code Quality," Mozilla.ai, 5 March 2026]. An academic pipeline uses an aggregator model to combine several distinct LLMs and reports up to +43.67% F1 over single-pass review [arXiv:2509.01494, 2025]. Both combine models without the measured independence our conditional requires.

Neither source, however, supports the claim that its system discards singleton findings. Star Chamber preserves them in a labelled tier, which is materially preferable to dropping them. The academic source does not describe its mechanism as singleton removal. We should not ascribe a common failure mechanism merely because both systems aggregate outputs.

Cursor's BugBot does document "majority voting to filter out bugs found during only one pass" [Cursor, "Building a Better Bugbot"]. It runs eight parallel passes on one model. Under §3, that is repeated sampling of a single stochastic witness, not a cross-vendor panel, and this conditional does not address it. We therefore have no named cross-vendor exemplar of the precise singleton-discarding mechanism against which the conditional was originally directed.

These distinctions also prevent an unwarranted practical conclusion. Precision-first triage under alert fatigue is a legitimate objective. Our criticism concerns treating correlated agreement as a truth procedure; it does not establish that these systems are wrong for their chosen operating points. Their claims should make the distinction clear.

The cost of retaining dissent

Our practice uses adjudication. Findings are not averaged. A blocking finding blocks, and a finding closes only when its raising reviewer withdraws it or a demonstration against the artifact settles it.

Four separately commissioned reviewers each identified the same weaknesses in this design. It grants a de-facto veto to any noisy, stubborn, or adversarial reviewer. It has no intrinsic bound on false positives, no stalemate rule for disputes without an oracle, and no cost model. A finding that cannot be outvoted creates an incentive to raise too many findings. Final authority also moves to a human adjudicator whose own reliability remains unmeasured (§11).

There is a further cost. A blocking finding is one sample from a stochastic witness. It carries sampling noise and must be established against the artifact. Treating it as a claim rather than a verdict is essential, but the work required must be included in the adjudication budget.

The method's viable operating region is therefore bounded: low-volume, high-stakes review, with oracle access and a limited panel. We withdraw any suggestion that it should replace high-volume triage. We also withdraw the circular claim that a lone dissenting witness preserves truth "by construction." The defensible statement is that under correlated majorities, only a procedure that retains and interrogates dissent can recover minority-true findings. Whether a given dissent is correct is what adjudication, or an oracle, must establish, at a cost that must be priced, not presumed.

One internal illustration helped motivate this design. A reviewer from another lineage reported fifteen issues, including one of high severity, after four same-lineage review rounds had missed them. This is a single unaudited record. It is an illustration, not evidence of a general advantage; the external research carries the argument of this section.

6. How a textual tradition emerges

So far, we have considered models as they stand. The paper's strongest claim concerns change through time.

Model-generated text entered training well before 2022: sequence-level distillation trained a student on its teacher's beam-search output [Kim & Rush, "Sequence-Level Knowledge Distillation," EMNLP 2016], back-translation trained translation models on machine-translated source sentences [Sennrich, Haddow & Birch, "Improving Neural Machine Translation Models with Monolingual Data," ACL 2016], and machine-translated text was already present in a web-scale pretraining corpus [Dodge et al., "Documenting Large Webtext Corpora," EMNLP 2021]. These examples establish that such transmission is not new; they do not measure its extent then or its growth since. Four channels matter.

  1. Direct distillation. A student learns from a chosen teacher's outputs, a deliberate and labelled practice established well before 2022 (Kim & Rush, 2016). Inheritance is experimentally demonstrated via watermark persistence into students (§4; Pan et al., 2025, with the underlying inheritance results attributed there to Sander et al. and Gu et al.).
  2. Unlabelled inclusion in collected text. Machine-generated writing enters the next corpus through ordinary collection. Estimates vary sharply. A keyword-frequency working paper puts the synthetic share near 30–40% [Spennemann, arXiv:2504.08755, 2025; not peer-reviewed]. A detector-based sample of 65,000 Common Crawl URLs reports a rise from about 2% in 2020 to roughly half in 2024–25 [Graphite/Originality.ai analysis, reported October 2025]. The reliability of AI detectors is itself disputed. Neither estimate should be treated as settled.
  3. Deliberate synthetic data. Laboratories disclose this use directly. Meta's Llama 3.1 model card reports "over 25M synthetically generated examples" in fine-tuning. Microsoft's Phi-3 report describes filtered web data combined with synthetic LLM-generated data [arXiv:2404.14219]. Anthropic's Claude 3 card includes "data we generate internally." Our description of curation as the only control on this channel is our own characterisation; none of these sources makes that claim.
  4. Diffusion through human writing. People adopt model phrasings and return them to the archive in their own text. Kobak and colleagues report an LLM-processing signal in 13.5% of 2024 biomedical abstracts, with higher levels in some subcorpora [Kobak, González-Márquez, Horvát & Lause, "Delving into LLM-assisted writing in biomedical publications through excess vocabulary," Science Advances, 2025, arXiv:2406.07016].

The last channel must be distinguished from the lexical example in §4. A shared methodology is a candidate explanation for why unrelated models overuse the same words. Their subsequent appearance in human writing concerns a different relationship: transmission from models to people. Evidence of that diffusion does not establish that the habit was inherited between models. Equally, a shared-tooling explanation for its origin does not invalidate a measurement of its later diffusion. One channel must not be allowed to serve as evidence for another.

Historical layers

Model generations deposit characteristic forms in the archive, and models trained on later crawls inherit them. That is dating by shared innovation in form; under the proposed rule, shared innovations are candidates for reconstructing relationships, and their evidential value remains to be tested (§7.1). In that sense, the archive acquires strata.

The qualification from §4 remains decisive. A deposited feature is only a candidate. Whether it is improbable enough under independent genesis to count as evidence of common ancestry is the measurement proposed in §7.1. The deposit is real. The reading of it waits on the instrument.

The loss of rare forms

Recursive training can erode a distribution's tails. Rare constructions disappear while probability concentrates on common ones [Shumailov, Shumaylov, Zhao, Papernot, Anderson & Gal, "AI models collapse when trained on recursively generated data," Nature 631:755, 2024; earlier "The Curse of Recursion," arXiv:2305.17493].

The training regime matters. In their language-model experiments, Gerstgrasser and colleagues report that replacing the original real data with each generation's synthetic data tends towards model collapse, whereas accumulating successive generations of synthetic data alongside the original real data avoids collapse in the settings studied [Gerstgrasser et al., "Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data," arXiv:2404.01413, 2024]. This qualification comes from that study, not from Shumailov and colleagues.

The philological comparison is with banalisation: repeated copying that favours easy or familiar forms progressively removes unusual ones. In the model case, the preference is built into the training objective rather than a scribe's fatigue or habit. Our argument is that the relevant training regime gives the purest measured instance of the mechanism philology inferred from its effects. Under that regime the harder-reading principle returns as a tendency in the tradition's development, not as a rule for choosing between two answers (§3). The same attraction towards probable forms that obstructs ancestry inference is a decay mechanism measured under pure replacement, removing rare forms through repeated transmission; whether it operates where synthetic material accumulates alongside real data, the regime in which Gerstgrasser and colleagues found collapse avoided, is not established.

The earliest surviving layer

The pre-2022 archive serves, in this argument, as the oldest witness available at scale; the examples earlier in this section show that it is not free of machine-generated text. McDonald introduced the analogy with pre-nuclear low-background steel on 5 December 2022, days after ChatGPT's release. Graham-Cumming's lowbackgroundsteel.ai subsequently catalogued relevant corpora.

The value of this earlier layer for baselining and dating rises monotonically as the later tradition develops, and depends on knowing its provenance. Conservation therefore becomes part of the research programme, rather than an incidental matter of storage: data provenance becomes archival science.

Deliberately planted features

A marker deliberately made improbable under independent origin provides a clearer case than a naturally shared habit. Maps and dictionaries have long used this principle. General Drafting Co. placed the fictitious hamlet Agloe in New York in the 1930s, forming its name as an anagram of its founders' initials. Christine Lindberg invented "esquivalience" for the New Oxford American Dictionary in 2001; it later appeared elsewhere. Both were planted to reveal copying.

An accidental ghost word is different. "Dord," a misread annotation that remained in Merriam-Webster's Second New International from 1934 to 1947, was not a deliberate copying detector. A Mountweazel is planted for that purpose. Watermarking and antidistillation fingerprinting use the same design principle: an engineered feature whose later recovery can reveal transmission. Their demonstrated inheritance is the cleanest existing proof that the descent channel is real (§4).

The possible security consequence

Code also enters training corpora. Consider a defective idiom produced in year N, committed to a public repository, and included in training for a model in year N+2. If its descendants retain the defect through that route, the vulnerability pattern has been inherited through the archive rather than independently regenerated.

The evidence does not yet establish that extension. Inheritance is demonstrated for engineered marks. Its extension to naturally occurring habits remains unresolved; its further extension to defective code idioms is an inference. It could be tested by retrospective tracing [verify: extension from the demonstrated channel to defect idioms; no direct study established here, check emerging literature].

If the inference holds, the channel is also an attack surface. Deliberately publishing a subtly defective idiom could poison an exemplar that later models learn from. This paper proposes tracing such transmission, never creating it. Planting defects in public corpora would attack the very resource the programme aims to conserve (§13).

Code is the best-conserved textual domain in the tradition. Version-control histories can preserve earlier states, hashes, timestamps, and signed records. An earlier state may be recovered by checkout and compared with a later tag. These records make code a strong candidate for conservation begun while the tradition is still forming, rather than reconstruction attempted after centuries of loss. Provenance, dating, and comparison with earlier witnesses are the proposed defences.

Three cautions remain throughout this section. The synthetic share of present crawls is disputed. Frontier laboratories filter their data aggressively. And this is an account of transmission structure, not of model intention.

Better models and more independent witnesses

It is reasonable to object that models are improving. More capable systems may help filter corpora, curate synthetic data, and produce fewer defects. Laboratories already use synthetic data in strong models, and the accumulation regime differs from the pure-replacement experiments that yield the strongest collapse results. We do not claim that models are generally getting worse.

Our concern is a different quantity. A performance benchmark asks how well a model answers. Comparison among witnesses also asks how differently they fail. A population may improve while becoming less useful as a collection of independent witnesses. Section 4's cross-sectional result bears on this distinction: more capable models more often agree on the same wrong answer.

That association can arise without copying; the study does not measure how much of it did. Corpus recursion is a second, distinct pressure on the same quantity. Its transmission channel is demonstrated, but its effect on the dispersion of model errors is argued here from mechanism, not reported as a measured trend. These two claims must remain separate.

Capability improvement alone therefore does not establish that §5's independence condition has been met. Section 7.3 proposes four responses. Preserving model vintages limits later inheritance, although it cannot remove convergent tendencies already shared when those models were sealed. Monitoring independence can detect correlation from either source without first separating the mechanisms. Increasing reliance on mechanical oracles shifts weight towards execution, tests, and proofs. Human specialists and formal methods are mandatory at the highest review tiers.

None of these claims to recover a defect that every witness silently inherits and no witness raises. That case remains beyond collation; §10 assigns it to interpretability and oracles (§11.7).

A thesis about change through time

The earlier reviewers were right to reject a direct genealogical reading of shared model errors. Our revised thesis concerns an emerging tradition. As outputs return to later corpora, transmission is creating material whose descent may become reconstructible; earlier, largely human-written material becomes harder to recover, while inherited features accumulate along demonstrated channels. Stemmatics exists only because traditions corrupt, and a tradition has begun to corrupt.

This is the paper's claim about the direction of development, not a completed measurement of the inherited fraction. We claim, without having measured it, that this fraction is small now and growing, which is precisely what makes the thesis time-indexed rather than wrong-then-right; it predicts that the stemmatic framework will become more applicable as the tradition deepens. More transmission does not by itself mean more recoverable descent: mixing of sources, and the loss of rare forms described above, can erase the features on which reconstruction depends. Section 11 states how that prediction could fail. The case for agentic stemmatics rests on this testable development, rather than on an identity between manuscripts and models.

7. The proposed investigations

The earlier draft was justly criticised for performing no distinctively stemmatic operation. This section sets out those operations as proposed designs with available ground truth. None is claimed as completed, and §13 states which component can be run today.

7.1 Reconstructing model relationships

The proposed experiment reconstructs relationships among models from their outputs and compares the result with documented lineages. Ground truth may include base-model and fine-tune families, teacher and distilled-student pairs, and successive checkpoints.

This is recensio: tracing relationships among extant witnesses. It does not reconstruct an answer that no witness supplies. Section 11.7 explains that limit.

The currently justified starting point is narrow: tracing planted marks and applying reference-based teacher tests to known candidate lineages [Rawat et al., arXiv:2607.09692], and, for pairs of models, screening for wrong answers that an unrelated background model does not usually give, as Stemma does (§10). This is legitimately a fingerprint-phylogeny benchmark. Collation of natural characters and grouping witnesses by significant shared habits require a further measurement: the improbability under independent origins of each habit, or of a set of habits assessed jointly with their dependence modelled. No study examined here supplies that measurement.

Tooling relationships form a second, separately labelled class. Tokeniser identity leaves measured, model-discriminating traces recoverable by black-box probing, including the under-trained-token failures discussed by Land and Bartolo (§4, with its pending full-text check). Active black-box probing still avoids access to weights. Two models sharing a vocabulary very probably share a laboratory, a codebase, or an open release. Such information concerns descent of tooling, not necessarily descent of text. These often documented relationships can help validate the reconstruction, but should not be presented as newly discovered textual ancestry. A reconstructed graph must distinguish the two kinds of edge.

What already exists

Earlier versions of this paper said that a complete reconstructed-and-validated model stemma appears to be an open research problem. As stated, that does not hold, and the sampled search behind it had missed the closest existing work. PhyloLM builds trees of language models from their outputs and checks them against disclosed lineage. For five models at the leaves of a Mistral-family tree whose relationships their creators disclosed, it recovers the known topology exactly; across 111 open and 45 closed models, it largely groups models by family [Yax, Oudeyer & Palminteri, "PhyloLM: Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks," ICLR 2025, arXiv:2404.04671]. One of its two sets of contexts comes from a code benchmark, and on that set the five-model tree again has the right structure, with its root misplaced.

PhyloLM compares the short continuations two models produce for the same contexts, and a shared continuation counts in proportion to how probable both models make it. Nothing in the score asks how often an unrelated model would produce the same continuation. PhyloLM draws its contexts from recent benchmark test sets to avoid familiar text on which every model would agree, which applies the right instinct to whole contexts but not to each shared feature. Its authors read the trees as showing at least training similarity, and they explain the mixing of three separately trained families by similar training sets: the distinction between common cause and descent drawn in §4.

For pairs of models a related test already exists: Stemma keeps probes on which the source model gives the same wrong answer in a majority of option orders and with a non-negative margin, and which one unrelated background model does not usually give; it ranks the survivors mainly by that margin and partly by how rarely the background model gives the answer, then measures how closely a suspect preserves the source's answers on the 40 probes it keeps. It outperforms a baseline that measures aggregate output agreement (§10).

What has been tested on code

An earlier version of this paper claimed, from abstract-level inspection, that neither attribution nor distillation-detection research included a code-domain experiment. That claim lasted one revision and was corrected at v1.22 after reviewers of two different model lineages, working separately, examined the full texts, including appendices.

For Sun et al., the conclusion stood. The evaluations used language-model output text only: UltraChat, FineWeb, Cosmopedia, LmsysChat, and WildChat. The attribution result was robust to paraphrase, but the reviewing lineages found no code-domain experiment.

For Rawat et al., the conclusion was false. The main evaluations concern mathematics and reasoning, but Appendix B.2 tests a change of domain: distillation uses mathematics prompts, while detection uses code-generation prompts from MBPP. The method reportedly remains accurate under that mismatch. A code-focused subject model elsewhere in the study is probed with mathematics prompts, which is not a code-domain evaluation either way.

Neither paper establishes an end-to-end experiment that both distils on code and detects on code. A mathematics-to-code probing transfer has been measured and held. A third paper, cited in §10, does both for an engineered marker: in the second version of antidistillation fingerprinting, a code-specialised teacher generates fingerprinted solutions to 674 MBPP problems, students are fine-tuned on them, and the fingerprint is detected in the fine-tuned students [Xu, Kirchenbauer, Savani, Trockman, Robey, Goldstein, Fang & Kolter, "Antidistillation Fingerprinting," arXiv:2602.03812v2, 2026, §4]. For features that nobody planted, the end-to-end code case remains untested in the studies examined here. The difference between a transfer and an end-to-end test, and between planted and unplanted features, matters both to feasibility and to the scope of what remains untested.

The review record behind this correction is internal and is not reproduced [unverifiable: internal record, not reproduced; permanent, not pending]. The underlying papers and named sections are public and can be examined without our record. An earlier draft omitted the marker because the limitation could never be discharged. That was a mistake: permanent inaccessibility strengthens the need for disclosure. We distinguish it from a pending verification rather than leave it unmarked.

What this experiment adds

The experiment proposed here must therefore improve on PhyloLM, not fill a gap PhyloLM leaves open. The difference lies in how similarity is scored. The test proposed here is whether scoring shared features, singly or jointly, by their improbability under independent origin recovers documented lineage better than raw similarity does, with PhyloLM as the baseline.

The proposal also extends Stemma's approach from pairs to a reconstruction over many models, from one background model to a population, from a background reproduction rate to an estimate of probability under independent origin, and to teacher and distilled-student pairs, where descent runs through text rather than weights. Scoring against a population is not new in itself: ErrorTrace already discounts a family's shared errors by the error rates of other model families, though for attribution to known families and without an independent-origin probability (§10).

That score needs a reference population in which similarity arises without descent. One candidate is the PolyPythias suite, which trains the same architecture on the same data with ten random seeds at each of five sizes, from 14 million to 410 million parameters, the seed setting initialisation and batch composition [van der Wal et al., "PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs," ICLR 2025, arXiv:2503.09543]. A habit shared across its seeds arises without any run learning from another's output; what the runs share includes their training data, hyperparameters, codebase, and architecture: common causes in the sense of §4. Nothing in this paper measures rates from it, and rates measured there need not transfer to models built differently; that transfer would itself have to be measured.

What remains provisional is therefore narrower. Trees built from output similarity exist, and so do provenance testing that compares a suspected child's agreement with each of its candidate parents against its agreement with unrelated control models, pairwise testing that keeps the wrong-answer probes an unrelated background model does not usually reproduce, and family attribution that discounts each shared error by how often other model families make it (§10). We have found no reconstruction over many models in which each shared output feature is scored against its probability under independent origin and the result is validated against documented lineage [verify: absence claim based on a sampled search, with scope declared].

Failures of reconstruction would also be informative. Multi-teacher distillation and mixtures of synthetic sources are forms of mixed transmission. The experiment should measure where the reconstruction breaks. We judge its evidential value high relative to its cost. It requires no weights access for collecting the subject models' outputs, and a negative result would materially weaken the thesis.

How the literature search was done

The absence claim above rests on a sampled search in three passes: an original search, which missed PhyloLM; a first pass on 24 September 2026 over a general web index and the arXiv record, which found PhyloLM but missed Stemma; and a second pass the same day over a general web index, the arXiv record and Crossref, with each reported work checked against its primary record and a forward-citation check on PhyloLM through Semantic Scholar, which found Stemma, Gallagher et al., Bruckner and TokenPrint. A later check against primary sources, made after the seat review of this revision had converged, found ErrorTrace, which Stemma cites and which the second pass had missed. No pass searched the ACL Anthology, and none covered humanities venues or literature outside English. Each of the three passes missed a close work, so the claim should be held loosely. It must not be confused with the full-text examination behind the code-domain correction. The four seats of the first review round of v1.33 were also asked for counterexamples from memory; none offered one, and three said they were answering without a search [internal round records, cited but not reproduced].

7.2 Measuring reviewer independence

We propose tracking error correlation and distances between reviewers' idiosyncrasy profiles, separately for different classes of question and over time. This would replace an assumption about vendors with measurements of the reviewers actually used.

The reviewers' finding stands: our earlier response-envelope identity check measures product identity, not error covariance. Knowing who answered does not establish how independently they answered. The proposed metric is not yet built. Its errors would require the same scrutiny as those of any other instrument (§8), and measurement would narrow rather than abolish the problem of a panel assessing its own independence (§11).

7.3 Preserving earlier witnesses

Frozen model weights are sealed witnesses. A model whose weights were fixed in year N cannot acquire year-N+1 contamination through later training. This permits an archive of open-weight vintages. Third parties cannot archive the weights of closed models, so only open-weight vintages can serve as sealed witnesses that answer new probes. Dated observations of closed endpoints can still be preserved: requests, responses, configuration and whatever version labels the provider discloses. HELM already releases its raw prompts and completions, with the times at which models were queried, for open, limited-access and closed models [Liang et al., "Holistic Evaluation of Language Models," TMLR 2023]. Such observations record what an endpoint returned; they cannot be re-queried once it is retired, and they cannot certify which weights produced them.

Such an archive would support three instruments:

For operating panels, we adopt a four-point degradation clause:

  1. Include witnesses from different training periods.
  2. Monitor the independence metric in §7.2 instead of assuming independence.
  3. Allow the fraction of review weight resting on mechanical oracles only to increase. Here mechanical oracles means execution, proofs, and tests. The normative-text consultation admitted in §3 is deliberately excluded: reading more documents is not the migration this commitment intends.
  4. Require human specialists and formal methods at the highest tiers as resources outside the machine textual tradition.

A method that measures its own decay can manage it. The proposal does not itself establish that the instruments will work, or that the response will be sufficient.

8. What our instruments got wrong

Our practice keeps a catalogue of recurring verification failures. Most concern false absence: a tool or operator reports that something is missing when it is present.

The catalogue includes quiet success mistaken for absence; matches discarded by the reader's filter; summary output mistaken for actual hits; a control that exercises a different command from the one supporting the claim; a mismatch between the kind of material tested by the control and the kind being searched; patterns that do not match the corpus's actual form; and differences between shell or runtime dialects. The control-to-corpus mismatch is an open family: the possible skip rules of arbitrary searches cannot be exhaustively listed in advance. Each catalogue entry has reproduced instances.

A safeguard that failed its own purpose

We built a gate intended to reject "not found" claims unless they carried a known-positive control. A reviewer other than its author constructed six routes to a false "absent." In one, the logic was inverted: a command's quiet success was reported as absence precisely when the target was present.

The gate's own five-case self-test had missed all six. Its author had tuned one control to a string that could not trigger the defect. The tool was quarantined, and its defects were preserved as permanently failing tests.

This establishes a possibility in one examined instrument: a verification tool can embody the very failure it is supposed to detect. It motivates two engineering responses. Controls must be capable of failing, and graders must face data that their authors did not construct. We do not infer a prevalence rate from the case.

The same class of error recurred during preparation of this paper. An operator claimed a document "does not exist" after searching a local directory, although the document was available on the public web. The catalogue's author had committed its own control-to-corpus mismatch. Absence claims here should therefore state the scope searched. An omission of that scope is another instance to correct, not an exception to the rule.

A corrected mutation result

An author-selected mutation set reported "5/5 killed, 100%." An AST-enumerated run on the same module produced 281 mutants. Of these, 143 survived, or 50.9%; 138 were killed, giving a kill rate of 49.1%.

The previous draft attached the kill percentage to the surviving count. Two reviewers caught the error separately. The underlying case illustrates how a generator-selected metric can encode the generator's own judgement in executable form. It is one artifact-backed observation, with no generalisation from n=1.

Two further events concern a sibling instrument, the pre-build protocol's frozen experiment commitments. They should remain distinct.

First, a critique asserted, among other things, that the instrument's central novelty claim was false. Primary-source inspection of the SARIF specification and two public issue records, sarif-spec#120 and in-toto/attestation#77, substantially settled that question in one pass, with edge corrections. The critique's three commissioned investigations had asserted the conclusion without performing that settlement. We did not measure their correlation and do not claim a value for it.

The sources also perform different roles. The specification is an oracle in §3's narrow sense: it establishes what the normative text says. The issue trackers are observations of filed public records. Neither establishes that a claim about the world is true. Together, however, they outperform three agents of one commissioned run, however convergent, on this bounded textual question.

Second, the critique was submitted to three external model lineages. They received byte-identical briefs, remained mutually blind, and our prior recommendation was deliberately withheld so no seat could inherit it. Two rejected that recommendation on textual grounds without seeing one another's answers. The third accepted it only as an explicitly declared extension of the governing clause [internal record of the sibling instrument, cited but not reproduced]. Section 9 explains the resulting decision and the limits of what the agreement established.

The verification pass also carried four minor inaccuracies in its own work, found by that same pass before use and not concealed: an undercounted enumeration, an inverted claim about a field's optionality, an unconfirmed attribution, and an uncounted variant. For that reason, our description remains "substantially settled, with edge corrections."

These are two dated instances, filed rather than hypothetical, of the proposed discipline in application. Where a suitable oracle existed, it was used. Agreement was considered in light of what the brief disclosed and withheld. The work reversed a conclusion reached in two prior internal passes. But the record concerns a different artifact and comes from one self-observed commercial practice (§11.10). It does not establish this paper's thesis or the method's effectiveness.

Other internal observations previously offered with load-bearing intent receive the same treatment. Rerun label instability, same-family blind-spot anecdotes, and capability-level claims about self-clearing are hereby downgraded to unaudited motivations. Their provenance entails selection, incentive, and reproducibility limits. Where external research covers the ground, that research carries the argument (§§4–5).

9. The comparison that was not run

An earlier draft said that "the experiment is specified." That was false at the time. The later record contains frozen commitments, but the instrument was discontinued before a registered run. This section records both developments.

The source is PBM-VALUES-V8, frozen on 16 August 2026 after eight adversarial review rounds. It recorded 30 findings: 29 closed by the reviewer who raised them, and one frozen as a declared residual. None closed by vote. Named files document the v7-to-v8 change, including one amendment made under the preceding round's "amend-then-freeze" disposition. This is an internal record, cited here rather than independently reproduced in full.

The stopping rule was declared by the commissioner, on the record and in advance. Round eight would be final regardless of its outcome. New findings would be recorded as residuals rather than prompt further amendments, so the record was not iterated until its reviewer tired.

The four frozen commitments

The following four passages retain the source paper's wording of the commitments. Paragraph breaks make them easier to read; the obligations and formula have not been rewritten. The expression "grading oracle" belongs to the frozen record. As the passage itself explains, it names a panel procedure, not an oracle in §3's sense.

Value 1 — module and arms. The module: a provenance tool, built twice from one frozen requirements sentence. Arm A: the practice's builder under the full protocol, with a two-seat review back-end. Arm B: a strong rule-governed baseline builder of a different vendor lineage, its transcript content an enumerated set (the requirements sentence verbatim plus a spec-compliance token). The record declares its own non-neutrality: the module is the protocol ecosystem's own wanted artifact, tolerable only because the case study decides nothing (Value 4).

A control envelope rule is frozen — the arms' shared system context is either empty or exactly one fixed string recorded in the pre-registration before either arm runs; content outside the frozen set voids the arm, and a void counts as a did-not-complete.

The fixed string itself is deferred to the pre-registration and is not yet recorded — we state this so the rule is not mistaken for its satisfaction.

A spec-compliance check is frozen by design: written before either arm starts, derived from the frozen requirements sentence and nothing else, output a bare PASS/FAIL token, identical for both arms, retries symmetric by construction.

Metering is frozen: metered API spend in US$ from driver response envelopes; Arm B is done at spec-pass, Arm A at spec-pass after its final review close, with all review costs inside the meter; hard cap US$75 per arm, exceeded-without-done counting as did-not-complete; token counts and wall-clock reported alongside, unthresholded.

Value 2 — the grading procedure (called "the grading oracle" in the frozen record; the name is that record's, not §3's, and it denotes a panel procedure rather than an oracle in this paper's operational sense — flagged so the two senses are not read as one). Panel-confirmed defects of severity ≥ MATERIAL only: blind finder raise, a written concrete failure scenario, and mandatory execution-candidacy with monotonic showing — a scenario that executes and shows is confirmed by artifact alone; failure to show is not-confirmed-by-execution; only a documented cannot-script determination reaches adjudication.

The practice's own instrument-defect catalogue is excluded from scoring, so treatment material cannot score itself.

Deduplication is frozen in content: two findings are one defect iff same file and same root failure condition, applied mechanically and arm-blind by the finders jointly; disputed clusters count as one defect and never route to the adjudicator.

Ledger-close is a frozen three-condition event: the stopping rule has fired, every cannot-script adjudication is filed, and every submitted script has been executed — close cannot occur while any submitted script is unexecuted.

Value 3 — roles. Finder seats named in the record with recusal rationale: a Moonshot-lineage seat and a Gemini-lineage seat, neither previously involved; the design's own reviewer lineage and the Arm-B builder lineage are excluded from finding for cause.

A named recorder — a principal of the practice who is neither builder nor adjudicator (personal names in the record; withheld here pending consent) — executes scenarios, saves every run output, and derives counts by script, with the recorder's one un-mechanised three-way classification declared as a residual rather than hidden.

Dispute rules are frozen: only finder seats and the builder may flag; the adjudicator cannot open a dispute; a dispute cannot unconfirm an execution artifact.

The adjudicator — the commissioner, conflict declared — is restricted to rows the recorder determines cannot be scripted, with symmetric powers only (confirm solely by endorsing the scenario, reject solely by endorsing another party's written objection).

The builder grades nothing, clusters nothing, runs nothing, and tallies nothing.

Value 4 — decision structure. The case study is frozen as non-decisional: it prices (cost per arm, defect profile, feasibility, retry and did-not-complete counts) and decides nothing.

The registered run's skeleton is frozen: minimum four completed pairs from an eight-module pool to be listed in the pre-registration before any pair runs, comparable scope, the case-study pair not reused, no arm-crossing.

Halt dispositions are frozen, including the asymmetric tiebreak: three protocol-arm did-not-completes halt the run as a failure; three control-arm did-not-completes halt it as baseline infeasibility — neither keep nor kill; pool exhaustion files as incomplete.

The stopping rule bars interim keep/kill decisions and peeking adjustments; and a declared selection limit scopes any verdict to the class of modules completable under the cap.

The success threshold, as a formula, verbatim: keep iff Σdefects_protocol < Σdefects_control and Σcost_protocol ≤ 3 × Σcost_control, a tie on defects failing the protocol arm deliberately.

The revision limit and its unresolved questions

The protocol allowed at most one revision, with the failure filed first and re-registration with a new pre-committed threshold. A second failure required discontinuation, and results had to be filed either way. Amending frozen values was prohibited; an attempted amendment itself triggered the discontinuation clause.

Two governance flags, recorded against this clause by the protocol's own reviewers, remain unadjudicated. The text did not say whether Values 1–3 remained protected during a failure-triggered re-registration. It also retained the confirmed residual on the ledger-close event: late submissions could refill the work queue. A candidate repair had been recorded but was held behind the amendment prohibition rather than applied silently.

On 17 August 2026, the instrument was discontinued. It did not fail a registered run; no such run occurred. The critique described in §8 had been substantially settled, with edge corrections (§8's phrase, carried here deliberately rather than upgraded), and considered by three external lineages.

What the agreement did not decide

An earlier account treated the two-to-one majority as settling whether the revision clause referred to a failed run or a defective instrument. It then called this "governance" to exempt the decision from §5's condition. That was wrong.

The meaning of the clause is a claim about a text. It could in principle be tested against the drafting record or contradicted by a fourth reading. Agreement among lineages whose dependence was not measured does not settle it. Reading their reasons before counting their votes does not establish that their judgements are independent. We withdraw the implication that the majority established the clause's meaning. The third reading remains in the record, and the interpretive dispute remains open.

The decision to discontinue rested on a narrower ground. The governing rule was contested, an event visible in the split itself, not a finding about the clause. Continuing with a contested instrument cost more than its remaining value. That decision did not require either interpretation to be correct, and the commissioner accepted discontinuation on that basis.

The reviewers shared an artifact and commissioning frame, and plausibly the training-corpus hyparchetype §12 names. Those shared inputs are possible channels of dependence, and they limit the evidential weight of their agreement about meaning. They do not invalidate a decision that did not depend on that agreement being true. Section 5 records why we no longer offer a general theory separating these kinds of decision. This case illustrates an attempted application of the narrower discipline; it does not demonstrate its success.

What the reader cannot resolve here

The disputed discontinuation clause is described, not quoted verbatim. Its internal source is cited but not reproduced. The ambiguity is at least visible in the references to "the failure filed first" and "a second failure": neither specifies whether an accepted critique of the instrument's design counts as a failure. Nevertheless, a reader cannot independently adjudicate the dispute from this paper. Our characterisation must be treated as unaudited.

These commitments are a closed historical record, not an active experiment. A proposed successor uses a narrower kill-screen standard: externally validated ground truth, a frozen mechanical matching rule, and blind scoring by someone other than the author. That is a separate commissioning, not a revision of this one. The reviewing lineages converged on discontinuing the original; they did not converge on the successor's particular design, which is our response [PBM-DISCONTINUATION-RECORD-2026-08-17.md; the successor's registration record remains owed].

Even a successful result under the retired design would have licensed only the claim that "this builder-plus-protocol bundle was a better purchase than this baseline under these rules." It would not have established that discipline works in general, that the result transfers to other builders, or that the protocol outperforms unconstrained generation. The baseline itself was rule-governed.

The case study was non-decisional: it priced the exercise and could decide nothing. The multi-pair study required at least four completed pairs from an eight-module pool that had yet to be frozen. The record says the pool "does not exist yet; the skeleton binds its shape and size." Nine open residuals included a permanent lineage/discipline confound, partial blindness of the finder reviewers, and an expectancy leak from the case study to the registered run.

The four values corresponded to the protocol annex's four required commitments. Whether they had separately been lodged with the recorder as a registration act remains [verify]. We claim a freeze, not that lodging. The protocol's own open finding D5 also preserved the transfer question: its strongest earlier evidence concerned review briefs, and transfer to generation had not been verified.

What the case record contributes

Retiring the comparison left unanswered whether this protocol produces better work than a strong baseline. It did not mean that the practice had produced no case material.

The practice records outside reviews of third-party code at pinned revisions. Findings distinguish reproduced from read. Reproduced means a deterministic test was run against unmodified production code at the stated revision. Read means source inspection supports the finding but no execution was performed. Records name controls, record positive dispositions for examined surfaces found sound, and carry a commitment to withdraw findings in writing if they do not survive maintainer scrutiny.

The publication limit is substantial. As checked on 23 August 2026, no case report or reproduction artifact, including a harness or pin instructions, was publicly available for a stranger to fetch and run. An earlier draft claimed that reports were public where disclosure had completed. No supporting pointer could be supplied because none had been published. Third-party disclosure timing partly constrained publication, but it did not change what a reader could verify.

The internal record establishes two narrow claims: that the defect shapes this programme names occur in production systems rather than only in its taxonomy, and that the practice can produce deterministic reproductions rather than assertions. Until public artifacts are available, both rest on unaudited records subject to §11.10.

The records also include nulls. One records a positive disposition for examined surfaces. Another discloses that findings from an engagement had already been patched upstream, after the practice's novelty gate caught the fact. These are offered as illustrations per §8's discipline: not as evidence, and not under "establishes".

A case series cannot establish that the multi-vendor review apparatus caused its successes rather than competent researchers using good tools. It lacks that counterfactual. Nor can it establish a detection rate, because missed defects are absent from the record. Comparative value against a strong in-house team remains the question the discontinued study did not answer.

Engineering crafts often accumulate worked cases before controlled comparisons. Those cases can later supply candidate ground truth for a more systematic science (§10). That describes the kind of record we possess. It does not substitute for the comparison we have not performed.

Several neighbouring fields address parts of the same problem. The relationships matter both to the proposed contribution and to its limits.

Model phylogeny, fingerprinting, and provenance

The work nearest to §7.1 can be sorted by the access it needs and the answer it returns. Several methods build trees over many models from their outputs, and some also or instead from their weights. PhyloLM scores agreement on short continuations (§7.1). LLM DNA embeds models' answers to benchmark prompts, checks the resulting representations against the model relationships recorded on Hugging Face, and builds a tree from them [Wu, Zhao, Wang, Guo, Wang & He, "LLM DNA: Tracing Model Evolution via Functional Representations," ICLR 2026, arXiv:2509.24496]. Gallagher and colleagues treat weights as genotype and outputs as phenotype, and recover the topology of a known training tree in a controlled experiment [Gallagher, Rallapalli, Brooks, Loughin, Sezgin & Yurko, "Analysis and Explainability of LLMs Via Evolutionary Methods," arXiv:2605.02930, 2026]. PhyloLM and LLM DNA summarise a pair of models by one aggregate similarity or distance; Gallagher's abstract does not say how it scores a pair.

Provenance testing asks a narrower question: was one model derived from another? Nikolic, Baluta and Saxena measure how often a suspected child and each candidate parent give the same first token across many prompts, test whether the best-matching candidate agrees with the child more often than the other candidates and a set of unrelated control models do, and report 90–95% precision and 80–90% recall on benchmarks of more than 600 models ["Model Provenance Testing for Large Language Models," arXiv:2502.00706]. That is the right kind of null, a population of unrelated models, applied to aggregate agreement with one suspected child at a time rather than to each shared feature. The authors also note that controls trained on very similar data raise the baseline and hide real derivation, which is the common-cause problem of §4 in another form.

Kuditipudi and colleagues test derivation from a model whose training order is known, using the tendency of models to memorise data seen later in training [Kuditipudi, Huang, Zhu, Yang, Potts & Liang, "Blackbox Model Provenance via Palimpsestic Membership Inference," arXiv:2510.19796].

One provenance test operationalises a feature-level version of the significance condition of §4. Stemma, whose authors take its name from stemmatics and set out the stemmatic analogy behind its choice of probes, keeps multiple-choice probes on which the source model gives the same wrong answer in a majority of option orders and with a non-negative margin, and which an unrelated background model does not usually give; it ranks the survivors mainly by that margin and partly by how rarely the background model gives the answer, then measures how closely a suspect preserves the source's answers on the 40 probes it keeps.

Across 770 pairs of source and suspect models drawn from 56 public checkpoints it reports an AUC of 0.967, ahead of four black-box baselines that include the aggregate first-token agreement of the provenance tester above [Zhang, Safronov & Martin, "Stemma: Induced Decision Regions Reveal LLM Provenance," arXiv:2607.25880, 2026]. It was posted on 28 July 2026, before this paper's first deposit on 20 August 2026. Its result is consistent with that condition: wrong answers that are uncommon in the background model help distinguish related from unrelated models. It does not estimate what share of shared errors arise through convergence, and its background rate describes one model's answers, not a probability under independent training.

Four limits separate it from the proposal of §7.1:

  1. It tests one pair at a time and builds no tree.
  2. Its null is a single background model.
  3. The lineage it tests is derivation from a base checkpoint through weight operations (fine-tuning, adapters, merging, quantisation), including further training on a stronger teacher's outputs, where the relation to the teacher is not the one tested.
  4. It builds each source fingerprint from the source model's next-token log-probabilities over the answer labels, which its authors call white-box choice label scoring, although checking a suspect needs only generated answers. An output-only extension would have to replace that construction step.

An earlier black-box method approaches the same condition from the population side. ErrorTrace weights each item on which a model family errs by the accuracy-weighted share of the family's models that err there, and discounts it by how often each other family errs there; it then places a suspect model in the family whose error space its own errors match most closely.

Across 27 models from four families it reports 85.18% attribution accuracy when the suspect is not among the models used to build the error spaces, and it traces 34 fine-tuned, pruned and merged derivatives [Zang, Meng, Chen, Cong, Zha, Qi, Li & Guo, "ErrorTrace: A Black-Box Traceability Mechanism Based on Model Family Error Space," NeurIPS 2025]. It appeared before both Stemma, which cites it, and this paper's first deposit. Its authors read consistent, uncommon errors as often reflecting shared architectural biases, training regimes or optimisation strategies.

Three limits separate it from the proposal of §7.1:

  1. It assigns a suspect to one of a set of known families and builds no tree over models.
  2. Its discount is an observed error rate in other families, which can share an error through common causes, not a probability under independent origin.
  3. The relation it tests is membership of a declared family, which mixes descent with common causes such as a developer's shared data and procedures.

Identification asks less again. LLMmap names which of 42 known model versions sits behind an application after a handful of crafted queries [Pasquini, Kornaropoulos & Ateniese, "LLMmap: Fingerprinting for Large Language Models," USENIX Security 2025, arXiv:2407.15847]. Bruckner fingerprints 165 served models from the distribution of their one-token answers to trivial prompts [Bruckner, "One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions," arXiv:2607.10252, 2026].

Instructional fingerprinting plants a private key in a model, which places it with the engineered markers of §4 [Xu, Wang, Ma, Koh, Xiao & Chen, "Instructional Fingerprinting of Large Language Models," NAACL 2024, arXiv:2401.12255]. Antidistillation fingerprinting chooses the marked tokens so that a student trained on them absorbs the mark efficiently, and detects distillation with less loss of output quality than earlier watermarks, again an engineered marker [Xu, Kirchenbauer, Savani, Trockman, Robey, Goldstein, Fang & Kolter, "Antidistillation Fingerprinting," ICML 2026, arXiv:2602.03812].

With access to weights, lineage can be read more directly: Horwitz, Shul and Hoshen recover directed model trees from model weights [Horwitz, Shul & Hoshen, "Unsupervised Model Tree Heritage Recovery," ICLR 2025, arXiv:2405.18432], and Zhu and colleagues test whether two models were trained from independent random initialisations [Zhu, Ahmed, Kuditipudi & Liang, "Independence Tests for Language Models," arXiv:2502.12292]. With access to internal states, TokenPrint finds that models trained independently on identical data score as more alike than fine-tunes of a shared base, which measures the common-cause problem of §4 directly [Wu, Zhao & Chen, "TokenPrint: A Calibrated Token-Space Fingerprint for Language-Model Provenance," arXiv:2608.08139, 2026].

In industry, Cisco's AI Supply Chain Provenance Explorer catalogues almost 900 open models and grounds their relationships in static fingerprinting and run-time behavioural fingerprinting rather than relying solely on self-reported lineage [Cisco, "AI Supply Chain Provenance Explorer for Responsible AI Governance," Cisco Blogs, 30 July 2026]. This programme assumes output access only, so methods that need weights or internal states are points of comparison and possible sources of ground truth, not substitutes.

The proposal of §7.1 sits between the first two groups: a tree over many models, as PhyloLM and LLM DNA build, in which shared features, singly or jointly, are scored against an estimated probability under independent origin, using a reference population designed to measure common-cause similarity without descent. Stemma supplies a pairwise precedent for feature-level scoring, against one background model. ErrorTrace supplies a precedent for scoring shared errors against the error rates of other model families, for attribution to a known family rather than reconstruction.

Assurance around untrusted models

AI Control develops protocols including trusted monitoring, trusted editing, and untrusted monitoring [Greenblatt, Shlegeris, Sachan & Roger, "AI Control: Improving Safety Despite Intentional Subversion," ICML 2024, arXiv:2312.06942]. It reaches the idea of assurance around an untrusted generator from alignment research. We approach it through software assurance and philology.

An earlier draft called this "decorrelated corroboration." We withdraw that inference. Both approaches belong to the shared discourse about untrusted language models after 2022. Their intellectual convergence is worth noting, but it is weak evidence of independent confirmation under our own rule (§4).

Generation constrained by formal verification is the oracle-space limit of this programme. Proof-checked synthesis and mathematical assurance programmes can supply machine-checked guarantees where suitable specifications exist. The UK's ARIA "Safeguarded AI," described here as a GBP 59M state programme pairing world models with machine-checked proofs, is one example. Our proposal to increase reliance on mechanical oracles deliberately moves towards that limit.

Interacting agents and shared memory

A recent survey of agentic reasoning maps systems in which several model agents coordinate through role assignment, communication protocols and shared memory, and in which agents are trained jointly rather than in isolation [Wei et al., "A Survey of Agentic Reasoning for Large Language Models: Towards Recursively Self-Improving and Collective Agents," arXiv:2601.12538v2, 2026, §§5.3.2 and 5.3.3]. Two of the mechanisms it describes bear on this programme. The first is memory that agents share: the survey distinguishes private from shared memory fragments and describes one design, Collaborative Memory, in which every entry carries immutable provenance (the source agent, the resources it accessed, a timestamp). The second is joint training, in which agents are optimised together; in one system the survey describes, three roles derived from a shared model backbone are trained jointly by reinforcement learning. The two act by different routes. A shared memory acts at run time: an output or a summary written by one agent can reach another without either model's weights changing, and this adds a candidate to the transmission channels of §6 that its taxonomy does not yet include. Joint training changes behaviour through optimisation. In the interacting systems the survey describes, another agent's output may contribute to an agent's training signal, a route adjacent to direct distillation, and a shared backbone or reward also supplies common causes in the sense of §4. Agreement following shared-memory access is therefore a candidate effect of transmission at run time; agreement following joint training may reflect transmission through training, common causes, or both. Neither shared access nor joint optimisation alone establishes the cause of an agreement. Where the evidence permits, transmission at run time, transmission through training, inherited weights (§6), shared external causes (§4) and independent convergence (§4) must be distinguished, allowing that more than one may contribute to the same agreement. A shared memory can make separate agents repeat one originating claim, and their agreement then carries a different evidential weight from independent rediscovery. We record both routes as candidates, not as measured channels. The survey maps these mechanisms; its own section on open problems, including the governance of such systems (§8, §8.6), leaves their control unresolved, and it does not establish that their effects permit reconstruction of model ancestry. Nothing in this paper measures them. The operational consequence for a review panel is the one §7.2 and §11.6 already draw: record what each seat could access before its blind return, and treat role specialisation (proposer, critic, judge) as a division of function that establishes no independence of lineage.

The inheritance from deterministic scanners

Earlier attestation formats already address adjacent problems. SARIF, an OASIS Committee Specification in 2019 and an OASIS Standard in 2020, defines six values for the per-result kind enumeration. These include fail for a match, pass for a clean result, and notApplicable. It also supplies artifact roles and content hashes to describe scanned material.

Microsoft's BinSkim and the OSS Review Toolkit provide related production mechanisms. BinSkim's --kind filter accepts Fail, Pass, Review, Open, NotApplicable, and Informational for scanned binaries. ORT's PathExcludeReason records why a path was excluded, using nine reason codes including BUILD_TOOL_OF, TEST_OF, and DOCUMENTATION_OF. The two classify different things: result dispositions and exclusion reasons. Their structures are comparable, not identical. FOSSology has recorded per-file human clearing decisions since its open-sourcing in December 2007.

These are ancestors of the proposed work. A new bespoke schema would not erase that inheritance.

Deterministic scanners seem to offer a connection between configuration and coverage, but that connection needs care. Configuration shows intended coverage. Runtime failure, unsupported formats, silent exclusions, early termination, and caching can prevent intended work from being completed (§8). Configuration alone does not prove examination.

Human reviewers and model reviewers present a further difficulty. Their attention across a repository is not observable to the consumer of their reports. Whether a model can internally represent or reveal that attention remains the interpretability question left open in §1. We do not settle it here.

The in-toto maintainers have publicly asked for a general human-review predicate. Their attestation repository's issue #77, opened on 5 December 2021, names chaining review coverage across a history of changes as an unresolved design problem. The source paper reports the issue as open at its time of writing.

We place the proposed collation apparatus in that line of work. Its intended contribution is closure as a verifiable property, rather than an indication of progress. But no section here yet specifies the requisite verifier artifact. Section 7.2 proposes an independence metric; §9 records a ledger-close rule from a discontinued instrument. Neither is that verifier.

At present, the contribution is a requirement, not an implemented verifier. If built, the intended packaging would emit SARIF results and an in-toto predicate. Only the closure verifier and judgement predicate would be claimed as new. No proposal has been submitted. Disclosure and timing remain separate decisions [verify: pointer to a submission record if and when it exists].

The multi-model review systems described in §5 are the closest prior work at the level of mechanism and the direct subject of that section's conditional.

Interpretability and collation

Mechanistic interpretability is complementary to the proposed study of textual populations. The analogy is with physiology and epidemiology. One examines internal mechanisms; the other studies patterns across a population.

Weights-based interpretability faces access boundaries between laboratories. Practitioners describe its current coverage as a small part of the computation, but we found no published figure quantifying that share in a sampled search of the literature and laboratory write-ups. This was not an exhaustive search; the absence claim is scoped here per §8's own requirement rather than asserted flat. We have also encountered the claim that exhaustive feature extraction would cost more compute than training, but found no quantified support and do not rely on it.

Agentic stemmatics would instead use outputs and the public component inventory required by §4. It would cross vendor boundaries without access to weights. Its raw material grows as the textual tradition develops, although the instruments proposed in §7 remain unproven.

Four considerations motivate that approach to the output-warrant question. First, the tradition crosses institutional boundaries that individual laboratories cannot inspect internally. Second, its archive can grow even as increasing scale plausibly makes complete mechanistic understanding harder. Third, outsiders can in principle rerun a documented collation, whereas they cannot rerun a weights-based analysis on inaccessible models. Fourth, philology supplies a historical example of reconstruction supported by a behavioural model of copying, without an account of the scribe's internal mechanism. Whether that sufficiency transfers remains to be shown.

Epidemiology offers a familiar precedent. Snow removed the Broad Street pump handle in 1854 without knowing the organism; Koch isolated Vibrio cholerae in 1883, twenty-nine years later. Action grounded in population evidence preceded identification of the organism. This is a historical comparison, not a validation of the proposed instruments.

Interpretability is stronger in several domains. It may reveal a corruption inherited by every witness, where no dissent remains for collation to compare. It can investigate intent, deception, and sandbagging that may leave no cross-witness trace. It supports causal intervention and training-time repair. Prospectively, it may also support mechanism-level absence claims that witness sampling can never establish.

The relationship already runs both ways. Findings that a model's self-commentary diverges from its computation are interpretability results on which this practice relies [Turpin et al., 2023; Lanham et al., 2023]. We propose collation as inexpensive, continuous triage across the model population that helps direct more costly mechanistic investigation. We have not operated that division of labour at scale.

The records such a programme would create, including dated strata, provenance manifests, sealed vintages, and archives of answers by question class, could also supply validation material for later mechanistic research. Industrial practice has played that role before: Pasteur's fermentation studies arose from spoilage problems in the French wine and brewing industries [Études sur le Vin, 1866; Études sur la Bière, 1876]. The comparison identifies a possible relationship between practice and science. It does not remove the need to validate this programme.

11. Limits

The framework is bounded by the following limits. Some were established by review of the earlier draft; others follow from the corrections made in response.

11.1 Some questions have no single archetype

Judgements about code may depend on specifications, priorities, or different legitimate questions. Reconstruction applies only to questions that an appropriate oracle could settle in §3's operational sense. Elsewhere, the panel samples judgement. Calling that reconstruction would restore an overclaim already withdrawn.

11.2 An available oracle takes priority

Where a test, type-check, proof, execution, or consultation of normative text can settle the relevant question, witness comparison is subordinate to it. The framework concerns the substantial remainder, subject to each oracle's own scope.

11.3 Witnesses must be sampled

A model lineage is stochastic. It must be sampled to be characterised; a single verdict remains a sample. This burden has no philological precedent and is unresolvable at the sample level, which prevents the straightforward transfer of some classical operations, including strict elimination of redundant witnesses.

11.4 Verification relocates trust

Trust remains in the human adjudicator, the independence metric, the controls, and the oracles. The adjudicator's error rate is unmeasured; capture and fatigue are unmodelled; in our practice, the commercial conflict is declared.

The independence metric in §7.2 is proposed and unbuilt. Its predecessors failed in two ways: an early statistical test was formulated backwards, and a later envelope check measured product identity rather than error covariance. The author-built positive controls also include a documented case of author-tuning (§8). Asking who grades the graders moves the question up one level. It does not finally answer it.

11.5 Independence cannot be fully established from within

Bédier's statistical objection concerned stemma-counting specifically, and the transposition to panel independence is ours, not his. Measurement can narrow the uncertainty (§7.2); nothing closes the problem of a system assessing its own foundations.

11.6 Dependence may follow the task rather than the vendor

Prompts, tools, briefs, common evidence, and assigned roles can introduce dependence among reviewers across vendors; shape diversity may dominate brand diversity. Our review history offers a qualitative illustration: findings about questions named in the brief followed the commissioning frame. We account for this when weighing agreement, but have no measured effect size (§12).

11.7 Unanimous silence is invisible

The method examines findings that witnesses emit. It cannot recover a defect that no witness raises. Classical emendatio has no counterpart here. A corruption shared by all witnesses leaves no disagreement to compare; §10 assigns that case to interpretability and oracles.

11.8 Shared sources are an attack surface

Shared corpora can introduce dependence. If the proposed extension from marker inheritance to defect idioms holds, corpus poisoning could also let an attacker induce correlated failures (§6). An assurance method must disclose this possible vulnerability in its own foundations.

11.9 The operating region matters

Retaining dissent dominates only where true minority findings exist at meaningful rates. In low-prevalence, high-noise regimes it is a false-positive engine. Section 5 therefore limits the proposed adjudicative approach to a particular operating region; it does not recommend it universally.

11.10 The field record is selected and internal

The observations come from one commercial practice examining itself. Client confidentiality, publication bias, incentives, and limited reproducibility are structural features of the record. Section 8 therefore treats internal observations as motivations, while §9 distinguishes frozen commitments from performed comparisons and public artifacts from inaccessible records.

11.11 Publication may change the phenomenon

If the programme succeeds, future models train on its methods and artifacts; detectors become optimisation targets; the planted marks must be renewed. The future validity of the method must therefore be tested through time as well as across models.

11.12 We do not know the denominator

The catalogue in §8 lists detected defects. It cannot count the false absences that remain undetected. It is an inventory, not a bound on the possible failures.

The same applies to this paper's correction record. The homoplasy correction, the withdrawn aggregation criterion, the misattributed citation, the stale maintenance notes, and the verification-marker count wrong six revisions running are all observed catches. We do not know how many unwarranted claims remain.

A long record of corrections shows that a correction process operates. It is also compatible with many errors remaining in a document while the readily detectable ones are diligently recorded. The practice that produced both the paper and its record cannot distinguish those possibilities from within. We know that the process can report failure. We do not know how often it should have done so.

11.13 Refuting our doctrine does not validate our method

An earlier internal epistemology based on falsificationism failed its own panel review; only structural theses survived. The previous draft called this "the strongest evidence the method functions." That inference was circular: rejection of the foundation counted as rigour, while acceptance would also have counted as rigour.

We withdraw it. The episode shows consistency in applying the practice's procedures to its own doctrine and recording the result. It does not establish the method's effectiveness.

11.14 The central thesis can fail

The claim about development through time in §6 fails if, over successive model generations, any of the following occurs:

The experiment in §7.1 is intended to expose the thesis to these possibilities.

11.15 The amount of scepticism also needs justification

We disclose the commercial interest in the method's adoption, the selection of the field record, the adjudicator's role, and the inaccessible evidence. Disclosure identifies the conflict. It does not measure its consequences.

External literature supports the correlation problem (§§4–5), and §11.9 limits the proposed operating region. Nothing here measures whether the amount of adversarial review this practice sells is proportionate to the generators' actual error rates. Even the discontinued experiment could have supplied only a bounded comparative result, not that general measure.

A practice selling adversarial review has a structural interest in the belief that such review is necessary. The framework's proportionality therefore remains an empirical question under a declared conflict of interest.

12. The review of this paper

Version 0.1 was reviewed by four model lineages from distinct vendors. They received the same adversarial brief, with the author's self-assessment appendix removed. All returned MAJOR-REVISION; none returned SUBMITTABLE or REJECT.

The four filed records were processed into a 209-row inventory of findings, items requiring verification, and requirements for incorporation. It was archived alongside the draft. The revised paper records responses to that inventory. Closure over each row belongs to the reviewer who raised it, not to the author.

The drafting lineage differs from the explicitly identified reviewing lineages, but it may not be unrelated to all of them. If Anthropic's allegations (Competing interests) are correct, models serving as seats from the four named lineages may have inherited features from Claude through its outputs, a possible instance of the direct-distillation channel described in §6. Such inheritance could affect the evidential weight of their agreement with Claude-drafted text, but neither its presence in the particular reviewing models nor its effect on their errors is established here. Procedural separateness alone establishes no independence (§3). Anthropic's September report also alleges that Moonshot and DeepSeek silently relayed some customers' requests to Claude and returned Claude's answers as their own; it reports measurement windows between May and July 2026, not an end date, and this paper's records cannot show whether any seat return was affected. The two Anthropic documents do not accuse OpenAI, xAI or Google, the developers of the other explicitly identified reviewing seats, Codex, Grok and Gemini, of these campaigns.

How the agreement was weighed

Some convergence followed questions already named in the brief: the claim of structural identity, unsupported effectiveness, and the risk of desk rejection. Agreement on these points may partly reflect commissioning and must be weighed accordingly (§11.6).

Other agreements were not directly elicited by those questions. Two reviewers, working separately, identified the absence of any distinctively stemmatic operation. Two separately caught the same arithmetic mislabelling. Four separately converged on the risk that adjudication could be obstructed by persistent or noisy findings. These observations shaped §4, §7, and the treatment of adjudication's costs in §5.

Their evidential weight is stronger relative to other findings within this commissioning, not demonstrably independent in an absolute sense. The reviewers still share the artifact, the genre of reviewer, and plausibly the training-corpus hyparchetype (§11.5–6). We did not measure their independence.

The narrow claim supported by this record is that the method was applied to its own artifacts, with recorded consequences and real cost. It does not remove the author's control over commissioning, the brief, or integration of the responses. Review reduces self-certification; it does not abolish it (§11.4).

The missing history of attempts

A critique reviewed by three external lineages mutually blind to each other's output, against a different artifact of this practice, identified a further weakness. If a review can be re-prompted, re-contextualised, and rerun until it closes, the apparent closure measures the author's persistence, not the finding's resolution, unless the record captures every attempt. A defensible record must preserve every attempt, including unsuccessful ones, outside the author's sole control.

That criticism applies here with undiminished force, in a specific and checkable way. The filed final verdicts retain the bytes associated with row closure, but the historical record does not distinguish a first-pass MAJOR-REVISION from a fifth response to the same prompt. Attempt counts, discarded runs, and prompt changes were not captured. The gap concerns how the artifact was reached, not merely what the artifact now contains.

At v0.3, the paper said that an append-only invocation log beyond the author's sole control would apply from the next review round. Seven revisions followed without such logging. This was a failure to perform the promised safeguard, not merely a delayed implementation.

Every round from v0.1 through v1.1 remains permanently outside contemporaneous attempt-level capture. A minimal append-only, hash-chained invocation log began at v1.1 and could cover only later work. Initially, the author alone held it. Such a log can provide tamper evidence to a later holder of a copy, but it did not supply the independent custody originally promised.

The subsequent improvements must be read against that permanent historical gap.

What the Bitcoin anchor established

On 20 August 2026, the first log entry's hash was committed in a Bitcoin OP_RETURN output. The transaction was:

542748138c9889e745c99fbccd268ec16cafc3e53ebb6fe30b1ee44cdd0145e0

It was confirmed in block 963282 at 09:23:01 UTC. The payload was the ASCII string CPRLOG1: followed by the SHA-256 entry_hash of entry sequence 1.

A reader with the log can compare that entry with the published commitment. The anchor establishes that the particular hash existed, and was published, no later than the block. Later alteration of the corresponding entry would create a detectable mismatch.

It does not establish custody of the contents by an independent party. A hash is not a copy of the record. Nor does it establish completeness. At that stage, the log contained four entries but only the first was anchored. The other three were appended on 20 August from filed artifacts, after the review rounds they described. They were labelled retroactive. Their artifact hashes were genuine, but their capture was not contemporaneous with the events.

The improvement therefore concerned tamper evidence for one entry in a four-entry log. It did not supply the independently held, complete history originally envisaged.

What publication changed

Later on 20 August, the log was published at github.com/clvstra/agentic-stemmatics-closure-log with a standalone chain verifier and the anchor record. The source paper records signed commits and a branch rule prohibiting force-pushes and deletions, including by administrators. Enforcement was checked by attempting a history rewrite and observing the server reject it.

This made more of the record checkable. A reader could clone it, recompute the chain, compare entry 1 with the Bitcoin transaction, and inspect later commit history without requesting the contents from us.

The remaining dependencies are important. The repository is under the author's account. Its second holder is the practice's other principal. GitHub is a trusted service in the arrangement. These facts do not provide the independent custody originally required.

The narrower gain is a public history against which changes can be checked. Silent revision would require defeating that history rather than editing a private file. Whether this satisfies the finding is for the reviewer who raised it to decide. The source record marks it addressed but still open pending that disposition.

13. Claims and responsibilities

Agentic stemmatics proposes a way to study the relationships among machine-generated texts and the models that produce them. Its central inference is a corrected rule for lineage in model populations: ancestry is carried by significant shared idiosyncrasies, among them a shared error that is itself improbable under independent origin, and a feature, or a set of features assessed jointly, counts only when independent origin of that agreement would be improbable. The rule is proposed, not established, and its condition would not be sufficient: by itself it identifies no parent, direction or route (§4).

Engineered markers and reference-based teacher tests supply evidence for the rule in their own transmission settings. Stemma adds evidence for pairwise provenance within checkpoint families, from fingerprints selected by stability, margin and low reproduction by one background model (§10); it does not establish independent-origin probabilities for the answers it selects. ErrorTrace adds evidence for attributing a model to a declared family from errors that other families seldom make (§10); its discount is an observed rate in other families, and family membership mixes descent with common causes.

No study examined here establishes the population-level independent-origin probabilities needed to use naturally occurring habits in a reconstruction over many models. Their use in stemmatics remains a research programme, not a result.

The paper also advances a conditional caution about aggregation: where reviewer independence is unmeasured or failing, agreement cannot be counted as independent corroboration. External measurement supports the correlation on which the caution rests; the loss of correct minority findings is our inference, not a measured rate (§5). The costs of our own alternative, adjudication, are stated alongside the risks of voting.

It argues that recursive transmission is creating a textual tradition through measured channels, and predicts, without having measured it, that this will make the stemmatic frame progressively more exact rather than presently identical; whether inherited features stay distinctive enough to reconstruct descent is the open question of §7.1 and §11.14. It proposes a validation experiment against known model relationships (§7.1). Only its planted-mark component, which would replicate existing results, can be run today; the natural-feature comparison with PhyloLM still requires a fixed feature set, reference population, estimator, reconstruction method, held-out split and success metric. It also specifies conditions under which its central thesis would fail (§11.14). The value of the proposed conservation programme depends on that thesis.

The constructive claim is one of possibility. Machine-generated work can have sufficient warrant for consequential use when the surrounding process supplies it (§1). This paper does not establish that our apparatus, or any particular apparatus, has achieved that standard.

We do not claim that the method produces better code. The controlled comparison was discontinued before execution (§9). We do not recommend adjudication outside its stated operating region, infer rates from internal anecdotes, or treat this paper's survival of review as evidence of its truth. Surviving collation that its author did not commission can add warrant to a claim; it does not make the claim a proof.

Ethical limits

The proposed study of corpus inheritance must trace existing transmission, never seed defects into public corpora. Deliberate planting would attack a shared resource that this programme seeks to conserve.

Client-derived observations remain confidential and are identified as unaudited where used. The author's commercial interest in adoption is a standing reason to distrust the unexternalised parts of this record. The argument therefore seeks support in external literature, frozen quoted commitments, and proposed experiments, of which only the planted-mark component can be run today. Whether that support is sufficient is for readers and reviewers to judge.

Outstanding verification

The source paper records a citation pass and distinguishes pending checks from permanent limits on access. Nine pending obligations remain in this revision:

  1. Philological definitions against standard handbooks (§2).
  2. The directionality of scribal errors (§3).
  3. The usus scribendi candidate anchors named at v1.8 (§4).
  4. Full-text confirmation of the under-trained-token study's scope and discriminatory power (§4).
  5. Grounding for the persistence of mode preference under temperature, instruction tuning, and preference tuning (§4).
  6. The extension from demonstrated inheritance to defect idioms (§6).
  7. The sampled-search absence claim, narrowed at v1.33, concerning a reconstruction over many models from output features scored against independent origin and validated against documented lineage (§7.1).
  8. Whether the frozen commitments were separately lodged as a registration act (§9).
  9. A submission-record pointer if the proposed interoperability contribution is submitted (§10).

The permanently inaccessible review record in §7.1 is a different kind of limitation. Its [unverifiable] marker does not enter the pending inventory and cannot be discharged by a future check of that private record on the reader's behalf.

Why the inventory itself required correction

The source paper's history of these markers is part of its evidence about instrument failure. In v1.22, two obligations were discharged. Full-text reading falsified the earlier claim about experimental coverage of code, and checking publication found that no case artifacts had been made public. Both outcomes became statements in the body. A check that disproves a claim discharges an obligation just as a confirming check does.

Version 1.24 added the under-trained-token obligation, taking the inventory from seven to eight. Review then required the mode-persistence obligation, bringing it to nine in v1.25. Those changes had to be reflected in the inventory as well as in the text.

Across six consecutive revisions, the count had been wrong through four identified mechanisms:

A fifth possible failure was avoided: discharging markers while leaving the total unchanged. Recounting as part of the edit prevented that mistake before a later reader found it. The lesson is procedural. Whenever the obligations change, the inventory must be reconciled with them. The count is not re-asserted as settled; it is stated as the product of the most careful method tried so far, and it is not a certificate that the underlying claims are true.

References

Entries carry caps on what may be assumed of them, and no positive verification mark. An earlier draft of this section tagged every entry the author had checked against a primary source. Two external reviewers, asked separately, both recommended removing that tag: a status field only its author can populate is indistinguishable to a reader from a false one, and placing it beside specific figures implied a content check the artifact does not substantiate. The exhibit was inside this very section — a FOSSology date carried the tag and still contradicted the body text by a year. The marks that remain are limits rather than credentials. [SEC] = established from secondary or reference sources only. [PP] = preprint, not peer-reviewed at time of writing. [REPORTED] = rests on press or vendor reporting this paper did not independently audit. An unmarked entry is an ordinary citation, and claims nothing about who checked it.

Correlated error, attribution, and distillation

Anthropic. (2026, 23 February). Detecting and preventing distillation attacks. anthropic.com/news/detecting-and-preventing-distillation-attacks. — an interested party's public allegation, cited in Competing interests and §12 for what it alleges, not as established. (Added v1.33.)

Anthropic. (2026, 10 September). Detecting and countering misuse of AI: September 2026. anthropic.com/threat-intelligence-report-september-2026; section "Illicit distillation", pp. 143 ff. of the PDF. — as above; the distillation section read for v1.33. (Added v1.33.)

Kim, E., Garg, A., Peng, K., & Garg, N. (2025). Correlated Errors in Large Language Models. Proceedings of the 42nd ICML, PMLR 267:30038-30066. arXiv:2506.07962.

Land, S., & Bartolo, M. (2024). Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models. EMNLP 2024, 2024.emnlp-main.649. arXiv:2405.05417. — cited for the prevalence and output-observability of under-trained-token artifacts tied to tokenizer vocabulary; checked at abstract and venue level only at citation time.

Kim, D. (2026). Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles. arXiv:2607.20768. [PP]

Kohli, G. (2026). Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels. arXiv:2605.29800; also published by Apple Machine Learning Research (machinelearning.apple.com/research/correlated-llm-evaluation-panels, checked 24 Sep 2026; the paper gives the author's affiliation as Apple). [PP] — title block, abstract, panel description (§3.2) and the same-family versus cross-family analysis (§5.3) read at v1 for v1.33. (Added v1.33.)

Sun, M., Yin, Y., Xu, Z., Kolter, J. Z., & Liu, Z. (2025). Idiosyncrasies in Large Language Models. ICML 2025. arXiv:2502.12150. — five-way attribution reported at 97.1%.

Pan, L., Liu, A., Huang, S., Lu, Y., Hu, X., Wen, L., King, I., & Yu, P. S. (2025). Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? ACL 2025 (Main), 2025.acl-long.648. arXiv:2502.11598. — cited for the inheritance premise it measures, not for its robustness conclusion; that conclusion is that inheritance is removable by rule recovery, dilution or generation-time neutralisation, not by ordinary paraphrase (v1.28 restated §6 accordingly). (Misattributed to "Zhao, X." at v0.5-v1.0; corrected on separate GLM and direct-fetch verification, both against the ACL Anthology entry.)

Sander, T., Fernandez, P., Durmus, A. O., Douze, M., & Furon, T. (2024). Watermarking Makes Language Models Radioactive. NeurIPS 2024 (Spotlight). arXiv:2402.14904. — the inheritance result §6 rests on: with open-model access and no knowledge of which texts were used, detection at p < 10−5 when no more than 5% of the fine-tuning instructions carried the marker; text-only access needs about 10%. (Added v1.28.)

Gu, C., Li, X. L., Liang, P., & Hashimoto, T. (2024). On the Learnability of Watermarks for Language Models. ICLR 2024. arXiv:2312.04469. — which tested watermark configurations a student learns detectably by distillation, and which it learns weakly or only slowly within the tested training budgets; sampling-based students train only on watermarked teacher samples, logit-based students on unmarked text (OpenWebText) against the teacher's watermarked next-token distribution; dilution of teacher output with other text is not tested. (Added v1.28; descriptor corrected v1.33.)

Rawat, R., Chen, S., Anand, A., Duan, M., Rotsted, B., & Min, S. Reference-Based Distillation Detection in LLMs. arXiv:2607.09692. [PP] — reference-based membership inference; despite adjacency in earlier drafts of this paper, it concerns no watermarking.

Kim, Y., & Rush, A. M. (2016). Sequence-Level Knowledge Distillation. EMNLP 2016, D16-1139. arXiv:1606.07947. — cited in §6 for training a student on its teacher's beam-search output before 2022. (Added v1.35.)

Sennrich, R., Haddow, B., & Birch, A. (2016). Improving Neural Machine Translation Models with Monolingual Data. ACL 2016, P16-1009. arXiv:1511.06709. — back-translation; cited in §6. (Added v1.35.)

Model phylogeny, fingerprinting, and provenance

Yax, N., Oudeyer, P.-Y., & Palminteri, S. (2025). PhyloLM: Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks. ICLR 2025. arXiv:2404.04671. — trees from output similarity (Nei similarity over short continuations, Neighbour Joining); 111 open and 45 closed models; main text read at v5 for v1.33, including its own summary of the Appendix D code-set result, and Equation 1 checked in the HTML source; appendices not read. (Added v1.33.)

Wu, Z., Zhao, H., Wang, Z., Guo, J., Wang, Q., & He, B. (2026). LLM DNA: Tracing Model Evolution via Functional Representations. ICLR 2026 (Oral). arXiv:2509.24496. — method, validation and tree passages read at v3. (Added v1.33.)

Nikolic, I., Baluta, T., & Saxena, P. (2025). Model Provenance Testing for Large Language Models. arXiv:2502.00706. [PP] — method section read at v2: aggregate first-token agreement between a suspected child and its candidate parents, tested against competing candidates and unrelated control models with Holm-Bonferroni correction. (Added v1.33.)

Kuditipudi, R., Huang, J., Zhu, S., Yang, D., Potts, C., & Liang, P. (2025). Blackbox Model Provenance via Palimpsestic Membership Inference. arXiv:2510.19796. [PP] — checked at abstract level only. (Added v1.33.)

Pasquini, D., Kornaropoulos, E. M., & Ateniese, G. (2025). LLMmap: Fingerprinting for Large Language Models. 34th USENIX Security Symposium. arXiv:2407.15847. — checked at abstract level only. (Added v1.33.)

Xu, J., Wang, F., Ma, M. D., Koh, P. W., Xiao, C., & Chen, M. (2024). Instructional Fingerprinting of Large Language Models. NAACL 2024. arXiv:2401.12255. — checked at abstract level only. (Added v1.33.)

Horwitz, E., Shul, A., & Hoshen, Y. (2025). Unsupervised Model Tree Heritage Recovery. ICLR 2025. arXiv:2405.18432. — weight-based; checked at abstract level only. (Added v1.33.)

Zhu, S., Ahmed, A., Kuditipudi, R., & Liang, P. (2025). Independence Tests for Language Models. arXiv:2502.12292. [PP] — weight-based; checked at abstract level only. (Added v1.33.)

Xu, Y. E., Kirchenbauer, J., Savani, Y., Trockman, A., Robey, A., Goldstein, T., Fang, F., & Kolter, J. Z. (2026). Antidistillation Fingerprinting. ICML 2026. arXiv:2602.03812v2. — abstract read; the experimental set-up (§4), including the MBPP code experiment, read at v2 (submitted 15 May 2026) for §7.1. (Added v1.33.)

Cisco. (2026). AI Supply Chain Provenance Explorer for Responsible AI Governance. Cisco Blogs, 30 July 2026. [REPORTED] — vendor description of a public catalogue; the fingerprinting methods were not examined. (Added v1.33.)

Zhang, K., Safronov, V., & Martin, A. (2026). Stemma: Induced Decision Regions Reveal LLM Provenance. arXiv:2607.25880. [PP] — full text read at v1 for v1.33 (method, implementation, benchmark composition, baselines, Appendix K); fingerprint construction uses the source model's choice-label log-probabilities, verification only generated answers; posted 28 Jul 2026, before this paper's first deposit on 20 Aug 2026. (Added v1.33.)

Zang, C., Meng, X., Chen, W., Cong, T., Zha, Y., Qi, D., Li, Z., & Guo, S. (2025). ErrorTrace: A Black-Box Traceability Mechanism Based on Model Family Error Space. NeurIPS 2025. — abstract, introduction, method (§4.1) and experimental set-up (§5.1) read in the proceedings version for v1.33; cited in Stemma's related work. (Added v1.33.)

Gallagher, S. K., Rallapalli, S., Brooks, T., Loughin, C., Sezgin, M., & Yurko, R. (2026). Analysis and Explainability of LLMs Via Evolutionary Methods. arXiv:2605.02930. [PP] — checked at abstract level only. (Added v1.33.)

Bruckner, T. (2026). One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions. arXiv:2607.10252. [PP] — checked at abstract level only. (Added v1.33.)

Wu, Y., Zhao, S., & Chen, J. (2026). TokenPrint: A Calibrated Token-Space Fingerprint for Language-Model Provenance. arXiv:2608.08139. [PP] — checked at abstract level only. (Added v1.33.)

van der Wal, O., Lesci, P., Muller-Eberstein, M., Saphra, N., Schoelkopf, H., Zuidema, W., & Biderman, S. (2025). PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs. ICLR 2025. arXiv:2503.09543. — cited as a candidate reference population, not for any result; §2 read at v2 (every run uses the same data, hyperparameters and codebase; seeds vary parameter initialisation and batch composition). (Added v1.33.)

Liang, P., Bommasani, R., Lee, T., et al. (2023). Holistic Evaluation of Language Models. Transactions on Machine Learning Research. arXiv:2211.09110. — cited in §7.3 for its release of raw prompts and completions, with query times, for open, limited-access and closed models. (Added v1.35.)

Model collapse and corpus recursion

Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature 631, 755-759. doi:10.1038/s41586-024-07566-y. Earlier: The Curse of Recursion, arXiv:2305.17493 (2023).

Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. [PP] — the accumulation-vs-replacement result; not part of Shumailov et al.'s own claims.

Kobak, D., Gonzalez-Marquez, R., Horvat, E.-A., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances. doi:10.1126/sciadv.adt3813. arXiv:2406.07016.

Juzek, T. S., & Ward, Z. B. (2025). Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models. COLING 2025. arXiv:2412.11385 — tentatively implicates RLHF; leaves the causal mechanism open.

Spennemann, D. H. R. (2025). Delving into: the quantification of AI-generated content on the internet. arXiv:2504.08755. [PP]

Graphite / Originality.ai analysis of 65,000 Common Crawl URLs, 2020-2025, reported Oct 2025. [REPORTED] — detector-based; detector reliability is itself contested.

Meta, Llama 3.1 Model Card; Microsoft, Phi-3 Technical Report, arXiv:2404.14219; Anthropic, Claude 3 Model Card. — cited for disclosed synthetic-data use in training pipelines.

Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., & Gardner, M. (2021). Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. EMNLP 2021. arXiv:2104.08758. — cited in §6 for machine-translated text in a web-scale pretraining corpus. (Added v1.35.)

Reasoning faithfulness and interpretability

Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023. arXiv:2305.04388.

Lanham, T., et al. (2023). Measuring Faithfulness in Chain-of-Thought Reasoning. Anthropic.

AI control and formal assurance

Greenblatt, R., Shlegeris, B., Sachan, K., & Roger, F. (2024). AI Control: Improving Safety Despite Intentional Subversion. ICML 2024. arXiv:2312.06942. — source of trusted monitoring, trusted editing, and untrusted monitoring. The "untrusted advice" protocol is a later, separate fellowship writeup and is not part of this paper.

Wei, T., Li, T.-W., Liu, Z., Ning, X., et al. (2026). A Survey of Agentic Reasoning for Large Language Models: Towards Recursively Self-Improving and Collective Agents. arXiv:2601.12538v2 (v1 18 Jan 2026; v2 20 Sep 2026). [PP]; cited for §§5.3.2 and 5.3.3 (shared versus private multi-agent memory with per-entry provenance; joint multi-agent training) and for §8 (open problems, governance at §8.6); title, authors, versions and section headings checked against the arXiv record and its HTML rendering of v2 at citation time, not against the PDF; 29 authors, abbreviated.

ARIA (UK Advanced Research and Invention Agency). Safeguarded AI programme.

Attestation formats and the scanner stratum

OASIS. Static Analysis Results Interchange Format (SARIF) v2.1.0. Committee Specification 01, 23 July 2019; OASIS Standard, 27 March 2020. — result.kind is a six-value enum: notApplicable, pass, fail, review, open, informational.

sarif-spec issue #120, "Identify files that were scanned" (opened 2018-03-10, resolved into the artifact-role design).

in-toto/attestation issue #77, "Defining a generalized predicate format for 'human reviews' of artifacts" (opened 2021-12-05 by adityasaky; open at time of writing).

Microsoft. BinSkim, docs/UserGuide.md.

OSS Review Toolkit, PathExcludeReason (model/src/main/kotlin/config).

FOSSology, per-file clearing decisions since its open-sourcing, December 2007.

Multi-model review systems

Wilson, P. (2026). The Star Chamber: Multi-LLM Consensus for Code Quality. Mozilla.ai, 5 March 2026.

Cursor. "Building a Better Bugbot." Cursor Blog. Cited for the aggregation mechanism as documented in the source (eight-pass majority voting); a self-published resolution-rate figure elsewhere in that source is not cited here, was not audited, and does no work in this paper's argument.

Benchmarking and Studying the LLM-based Code Review. arXiv:2509.01494 (2025). [PP] — SWR-Bench; Multi-Agg and Self-Agg aggregation variants.

Philology and textual criticism

Maas, P. (1927). Textkritik. English: Textual Criticism, trans. B. Flower, Clarendon Press, 1958. [SEC] — source of Leitfehler.

Timpanaro, S. (1963). La genesi del metodo del Lachmann. English: The Genesis of Lachmann's Method, ed./trans. G. W. Most, University of Chicago Press, 2005. [SEC]

Bedier, J. (1928). On the manuscript tradition of the Lai de l'Ombre. [SEC] — 105 of 110 surveyed stemmata were two-branched.

Cerquiglini, B. (1989). Eloge de la variante. English: In Praise of the Variant, trans. B. Wing, Johns Hopkins University Press, 1999. [SEC]

Roelli, P. (ed.) (2020). Handbook of Stemmatology: History, Methodology, Digital Approaches. De Gruyter. [SEC]

Roos, T., & Heikkila, T. (2009). Evaluating methods for computer-assisted stemmatology using artificial benchmark data sets. [SEC]

Reynolds, L. D., & Wilson, N. G. (2013). Scribes and Scholars: A Guide to the Transmission of Greek and Latin Literature (4th ed.). Oxford University Press. Orig. 1968. [SEC] — canonical general account of scribal transmission; cited for that, not as a confirmed primary source for usus scribendi specifically. (Duxfield's 2019 digital-humanities project page, cited in earlier drafts, uses the term correctly but is not the primary source; replaced on review.)

Pasquali, G. (1934). Storia della tradizione e critica del testo. Le Monnier. [SEC] — named by a philological review of this draft as the more precise candidate anchor for usus scribendi as a technical concept; not yet confirmed against the primary text.

Bischoff, B. (1990). Latin Palaeography: Antiquity and the Middle Ages, trans. D. Ó Cróinín & D. Ganz. Cambridge University Press. Orig. German 1979, 2nd ed. 1986. [SEC] — named by the same review as an alternative candidate anchor for scribal-habit terminology; not yet confirmed.

Historical analogies

Snow, J. (1854), Broad Street cholera investigation; Koch, R. (1883), isolation of Vibrio cholerae. [SEC]

Pasteur, L. (1866). Etudes sur le Vin; (1876) Etudes sur la Biere. [SEC]

McDonald, K. (5 December 2022), first public use of the low-background-steel analogy for pre-AI-contamination data; Graham-Cumming, J., lowbackgroundsteel.ai. [SEC]

Agloe, New York — trap street placed by General Drafting Co. [SEC] — "esquivalience," coined by C. Lindberg for the New Oxford American Dictionary (2001) as a copyright trap. [SEC] — contrast "dord," an accidental ghost word in Merriam-Webster's Second New International (1934-1947). [SEC]

Still unresolved

The following claims in this draft remain uncited and are marked in text: the paper's own internal records (§8, §9, §12), which are cited but not reproduced; the extension of watermark inheritance to defect idioms (§6), explicitly an inference with no direct study; the §7.1 absence claim, narrowed at v1.33 after the earlier search missed PhyloLM, about a reconstruction over many models from output features scored against independent origin and validated against documented lineage, which rests on a sampled search with declared scope; the §7.1 two-lineage record behind the code-domain correction — internal, cited and not reproduced, carrying a permanent [unverifiable] marker; and the registration-lodging question in §9. These are named here rather than left for a reader to discover. One marker species besides [verify] appears in this draft: [unverifiable: ...], attached by the author to a claim whose supporting record is permanently internal. It never closes, belongs to no pending inventory — §13 counts pending verifications only — and its instances are listed here instead, so a reader auditing open obligations finds the permanent limitations beside the pending ones.

Appendix A: Revision history

A note on this build. The drafting copy of this paper carries a working-notes section at the end: an internal maintenance list of open review rows, owed work, and venue thinking, marked in the source as delete-before-submission. It is deliberately not published here. Several entries below refer to it as "the Notes"; those references point at that omitted material. Nothing in the argument, the References, or this history depends on it, and the open findings it tracked are also stated in the body where they bear on a claim. The entries below are reproduced unedited, which is why each carries the "NOT SUBMITTED" stamp it had when it was written: they are dated records of what was true at the time, not statements about this preprint.

Draft v1.37, 26 Sep 2026. Corrections to the record of this paper's review; no claim in the body changes, and one sentence is added to the search note in §7.1. The practice's in-house advisor, of the drafting lineage, read the round records of 24–25 Sep 2026, and a Kimi seat (Moonshot) reviewed that reading against the same records [internal records of both readings, cited but not reproduced]. This entry corrects four statements in the entries for v1.35 and v1.36 and records two omissions: shared findings in confirmation material and the loss of the first-round checksum record. (1) The v1.35 entry says that one return identified its harness as the tool named in that file rather than the tool used. Four Gemini returns reported ‘Harness: ZCode’. The collation identifies a Gemini self-report as wrong against the driver's log. Two quoted Kimi returns mention the same instructions file but identify their harness as Kimi Code CLI. These materials still do not establish that the file caused the identification. (2) The same entry says that returns from Kimi and Gemini seats in most rounds cite or report receiving that file. The advisor reports received vocabulary in 10 of 21 Kimi return files, with another 4 reporting an empty instructions block, and in 4 of 5 completed Gemini returns. Kimi reports reproducing these counts by searching for instruction-file vocabulary, reading hits in context and classifying receipt, empty-block reports and absence, with duplicates and probe/runlog files excluded from seat counts. On those reported counts, ‘most’ applies to Gemini returns and the pooled returns (14 of 26), but not to Kimi returns alone. These are counts of return files, not rounds. (3) The same entry says that a Codex global instructions file remained in place throughout the rounds, and its closing sentence says that the Codex seat ran with that file set aside in the fourth and fifth rounds. The supplied records document set-aside and restoration with hash checking for rounds four and five, the two review rounds of Introduction v0.19, the check of the library filing, and the deposit check. ‘Throughout the rounds’ is therefore incorrect. The records establish receipt of the global instructions for the probed invocation, not for every earlier run. The filed probe records distinguish two reports: one quotes the global instructions file's heading and opening sentence; the fourth-round runlog records another probe after the file was moved aside. That second response says, ‘No instruction block labelled AGENTS.md was provided in this session’, but also answers ‘(1) Yes’; the corresponding question is not reproduced here. The earlier entry's phrase ‘reported no such file’ therefore overstates the quoted response [probe records, cited but not reproduced]. The earlier entry's closing sentence, ‘set aside at the author's decision, after a probe reported no such file’, also reverses the recorded order: the runlog records the move before that probe. Both readings also report that no filed brief put the combined passage to a seat: the round-five brief excluded the outcome placeholder containing the closing account. (4) The v1.36 entry says that the Codex deposit seat had no network access. The seat reported unsuccessful browser fetches for the DOI URLs and deposit endpoints, but a successful site-page retrieval returning the v1.33 reading copy. Its confirmations were based on responses captured by the drafting assistant where identified, and the seat expressly said they did not certify current network state. The entry's unqualified account of access is therefore inaccurate; these records do not establish the retrieval architecture or the full extent of network access. (5) The v1.33 round-two brief listed all four first-round returns as available material. The v1.35 round-two row map presented the named seats' findings and instructed reviewers to close their own rows and treat the others as context. Both readings report this shared-findings design in subsequent confirmation rounds and place the four-seat catch of the Appendix item (2) residue in v1.33 round two. This design supplied the channel that §3 compares with copying from a second exemplar. The earlier entries describe byte-identical briefs and closure procedures but do not disclose the sharing of other seats' findings. These records do not establish that any seat copied another's conclusion or that sharing caused the convergence. (6) The checksum record of the first round of the v1.33 review was overwritten when a command chain run by the drafting assistant failed partway; its exact contents are lost [internal record, cited but not reproduced]. The custody note reports that the brief and single-file bundle still match the hashes recorded in the collation; for the other corpus files, unchanged status since serve time is not proven. The search note in §7.1 now adds that the seats of that first round were asked for counterexamples from memory and offered none. The entries below stand as written, as this appendix's practice requires; these corrections supersede them where they differ. This entry and the §7.1 sentence went to a Codex seat (OpenAI) and a Kimi seat in three rounds, each round under one brief: the Codex seat returned REOPEN, then CONFIRM-WITH-EDITS, then CONFIRM; the Kimi seat CONFIRM-WITH-EDITS twice, then CONFIRM. Replacement text a seat gave was applied verbatim; where both seats gave text for the same sentence, the Codex text was used; and text of the drafter's own beside a seat's text went back to both seats in the next round. By the third round each seat had closed every finding its lineage raised. The Codex seat ran each round with its tool's global instructions file set aside and restored, hash-checked [round returns and runlogs, 26 Sep 2026, cited but not reproduced]. These last four sentences were not put to a seat.

Draft v1.36, 25 Sep 2026. Housekeeping only. The front-matter line now records that v1.35 was deposited on 25 Sep 2026 as version DOI 10.5281/zenodo.22959463, with the deposited file's SHA-256; Zenodo reports its MD5 as baedd1a10e8b80b9f1bc582be3483b93. The deposit was checked through the Zenodo API and by two seats under the same brief. Kimi (Moonshot), from its own fetches and hashes, confirmed that the concept DOI and the version DOI resolve to the new record, the record's metadata, its single file (name, size and MD5), that file's byte identity with the PDF printed from v1.35, the five versions, and the live page the PDF was printed from. Codex (OpenAI), which had no network access and whose sandbox blocked hashing, confirmed the concept DOI, the metadata, the file listing and the versions from responses captured at the time. The version DOI resolved only once DataCite had registered it, about 100 minutes after publication: both seats first found it unregistered, and Kimi confirmed it resolving at 18:01 UTC [SEAT-DEPOSIT-FACTS-v1.35-KIMI and -CODEX, 25 Sep 2026, cited but not reproduced]. The record's version field and creator affiliation were corrected on Zenodo the same day, before the check. No other text changed.

Draft v1.35, 25 Sep 2026. Correction revision before deposit; claims changed. Two outside reviews of v1.33 were read against the paper and its sources: one recommending major revision, and one, known here only from its author's summary, a directed read recommending minor revision. The author reports both as reads by the Kimi lineage (Moonshot), one of the four lineages named in the allegations recorded in Competing interests; the first states no lineage in its own text. The drafting lineage checked each item against the paper and the primary sources, for (1) to (3) with an in-house advisor's reading of what the paper's own account of philology supports; that check is not independent of the drafting (§12) [internal records of the check and of the advisor's reading, cited but not reproduced]. The author reports that review seats received standing instructions not stated in the byte-identical briefs [internal record of the exposure, cited but not reproduced]. For the recorded v1.33–v1.35 rounds, the author reports that returns from Kimi and Gemini seats in most rounds, and Codex in three, cite or report receiving a workspace instructions file written for another lineage above their working directory. The author describes that file as procedural and as naming no claim of this paper, and reports that one return identified its harness as the tool named in that file rather than the tool used. These materials do not establish that the file caused that identification. The author also reports that a Codex global instructions file remained in place throughout the rounds and that a probe using the seats' invocation reported receiving it. This suggests possible overlap with the drafting assistant's owner context, but historical exposure across the Codex rounds and the extent of that overlap are not independently established here. The author further reports that Kimi returns from the last two rounds acknowledged receiving summaries of the owner's review skills and said those skills were not used. Those returns are not reproduced here; they do not establish an every-run loading claim here. Earlier rounds were not checked. Review records must distinguish the common brief from additional context actually supplied to each seat. Items found sound and material to the deposit are corrected here; the rest are recorded below for the next version. (1) The abstract and §§1, 3, 4 and 6 said that shared model errors chiefly reflect convergence, are predominantly polygenetic, or mostly arise independently. No cited source measures that share: the main study measures agreement between models when both are wrong, in a population that includes other developers' fine-tuned adaptations of base models, without a covariate for lineage. The text now says that shared errors are common and can arise without copying, and that what share did is unmeasured; the argument from philology's exclusion rule does not need that share. (2) §4 stated the central rule as an "if and only if", and §6 asked whether a feature is improbable enough "to prove descent". The rule is now stated as proposed, not established, and its condition as not sufficient: by itself it identifies no parent, direction or route. §6 now asks whether a feature counts as evidence of common ancestry. (3) §§4 and 13 required each feature to be improbable on its own; for models the condition may also be met by a set of features assessed jointly, with their dependence modelled; the existing methods of §§7.1 and 10 score features in aggregate. Statements of the rule in the abstract and §§1, 3, 4, 7.1, 10 and 13 now allow a set of features to be assessed jointly, and those in the abstract and §§1, 3, 4 and 13 call the rule proposed. (4) §13 described the conditional argument of §5 as an aggregation result grounded in external measurement; it is a conditional caution, and the loss of correct minority findings is marked as our inference. (5) §6 said that before about 2022 models contributed nothing back to what their successors would read; sequence-level distillation (Kim & Rush, 2016), back-translation (Sennrich, Haddow & Birch, 2016) and machine-translated text in a web-scale corpus (Dodge et al., 2021) show otherwise; the text now says such transmission is not new, without claiming a measured change in its extent, "now" is removed from the abstract's sentence about outputs returning to training corpora, "earlier uncontaminated material" becomes "earlier, largely human-written material", the pre-2022 archive is no longer described as preceding the contamination, and §6's claim that the inherited fraction is small now and growing, and the abstract's "is growing", are marked as unmeasured or argued, and §6 no longer says that shared innovations reconstruct relationships, only that they are candidates to be tested. (6) §§6 and 13 treated more transmission as more reconstructible descent; the text now says descent may become reconstructible, that mixing and the loss of rare forms can erase the features reconstruction needs, and that the direction stated in §13 is a prediction not yet measured. (7) §6 called the attraction towards probable forms a law of the tradition's development and its measured decay mechanism; the decay was measured under pure replacement, and the text now says so; §6 now reports Gerstgrasser and colleagues' language-model result (collapse avoided when real data are kept and synthetic data accumulate) in place of a bounded test error, which they prove in a simpler linear setting. (8) §4 attached the count of all 31,900 two-, three- and four-model subsets to three-model subsets and gave an after-the-fact rate without saying so; it now says the subsets overlap, and gives 9.98% of the 4,060 three-model subsets after the fact and 18.71% with held-out selection, and applies the author's "modest" and "configuration-dependent" to the association they describe. (9) §10 filed a shared memory and joint training together as transmission outside training data; they act by different routes, joint training can also create common causes, and the per-entry provenance the survey describes is now attributed to the one design it names, Collaborative Memory. (10) §7.3 said closed models cannot be archived by third parties; their weights cannot, but dated observations of their outputs can, as the completions HELM releases show. (11) §7.3 called perplexity exquisitely sensitive for dating; the probe is now stated with its controls, as a proposal to be tested. (12) §§7 and 13 called the validation experiment executable, and §§1, 3 and 13 said experiments were specified; the experiments are now outlined or proposed, only the planted-mark component can be run today, and §13 lists what the natural-feature comparison still requires. (13) The last sentence of §13 said a preproof becomes a proof by surviving collation it did not commission; it now says such collation can add warrant, not proof. (14) The abstract's account of the two dated events of §§8–9 now cites those sections, identifies the sibling instrument of §8 with the comparison of §9, summarises their source check and the recorded split with its separate ground for discontinuing, and uses their phrase "substantially settled, with edge corrections", which v1.34's "settles" overstated. (15) The Wei et al. reference gave 28 authors; the arXiv record lists 29. (16) The heading of the v1.33 entry below said 24 Sep 2026, although that entry records a round held on 25 Sep; it now says 24–25 Sep 2026. This is the only change to an earlier entry. New references: Kim & Rush (2016); Sennrich, Haddow & Birch (2016); Dodge et al. (2021); Liang et al. (2023). Editorial changes, from a presentation reading of v1.34 supplied by the author: the opening of §1 merges two paragraphs, removing a framing question whose examples repeated those of its first sentence; §7.1 gains three subheadings beside the existing one on code, separating what already exists, what the proposed experiment adds and how the literature search was done, with its paragraphs split by function and moved, with small linking and referential edits; in §10 the longer paragraphs are split by function, the final comparison sentence is divided into three, and the limits that separate Stemma and ErrorTrace from the proposal of §7.1 are set out as numbered lists; §13 separates its central claim from its evidence. The frozen commitments of §9 are unchanged. Recorded for the next version, not changed here: the unmeasured starting state of the historical transition, including "small now and growing" and an archival value said to rise monotonically; an estimator specification for the rule (reference population, feature selection, held-out evaluation, multiplicity and an "unidentified" outcome); the standing of the voting argument as a caution, the same-model exception, and whether measured pairwise independence determines when voting helps; separate tests for the extent of transmission, the persistence of source-specific features and the value to review; designs for the runtime-transmission and dating probes; a full protocol for the natural-feature experiment; a close teacher-tracing baseline the search missed (Wadhwa, Shaib, Amir & Wallace, Findings of ACL 2025); literature on same-family error and on agreement against correctness that the second review names for §§4–5, not checked here; an identifiability limit the second review reports from an appendix of TokenPrint, not checked here; the meaning of "independent" in "independent genesis"; and the length of the paper and of this appendix. v1.34 was prepared for deposit but was not deposited; this version replaces it as the version prepared for deposit, and the v1.34 entry below stands as written. Seat review of this revision: five rounds on 25 Sep 2026, to Codex (OpenAI), Kimi (Moonshot) and, in the first three, a Gemini-lineage seat (Google), under briefs byte-identical across reviewers, each later round going to the reviewers whose rows it answered, and from the third round with the full text of the draft; findings were corrected and put back to the reviewers that raised them, who closed them, with three exceptions. Codex's last two rows, on the exposure passage above, were answered by its own replacement text, applied verbatim except that "the editor" became "the author" and "in this corpus" became "here", and not put back. The Gemini seat's two remaining rows were answered, one by correcting a record error that row had introduced, and not put back, because the quota of its command-line tool was exhausted from the fourth round. The rows of the two outside reviews have not been returned to their raiser. In the fourth and fifth rounds the Codex seat ran with its tool's global instructions file set aside at the author's decision, after a probe reported no such file; its environment still supplied a catalogue of skills and agents that describes the practice, and the Kimi tool listed the owner's skill summaries, as both seats' returns report.

Draft v1.34, 25 Sep 2026. Deposit revision; no argument text changed. The author line and both competing-interest statements incorrectly identified the practice as "CounterProof Research Ltd": CounterProof Research is a practice of Clavestra Capital Limited (Malta, C 113987), and the author is the practice's director and co-founder. Those three places, the page header and the page footer now say so, and the author line gives that role in place of "Chief Strategy Officer". The four earlier deposits (v1.15, v1.25, v1.29 and v1.31) also contain that incorrect name and remain unchanged; this entry records the correction. The front-matter deposit line is rewritten so that it stays true once this version is deposited, instead of going stale as the entries for v1.26, v1.30 and v1.32 record of its earlier form. This revision was prepared for deposit as a PDF printed from this page, with the intended Zenodo metadata changes being correction of the author's name order and identification of the affiliation as CounterProof Research, a practice of Clavestra Capital Limited; the revision after the deposit records whether both were done.

Draft v1.33, 24–25 Sep 2026. Argument text revised after outside review; a site reading copy, not deposited. Five reviews of v1.32 by AI models of at least two vendors, each asked by the author without context, were read against the paper and against their sources. (1) §7.1 had said that a complete reconstructed-and-validated model stemma appears to be an open research problem. As stated, that does not hold: PhyloLM (Yax, Oudeyer and Palminteri, ICLR 2025) builds trees of models from output similarity and checks them against disclosed lineage, and the sampled search behind the claim had missed it. §7.1 now names PhyloLM as the baseline the proposed experiment must beat, states the difference (its similarity does not score a shared feature against what unrelated models would produce), names a candidate reference population for that score (PolyPythias), and narrows the absence claim, which keeps its [verify] marker; the §13 inventory entry and the References note on unresolved claims follow it, and the number of pending verifications is unchanged. (2) §10 gains a subsection on model phylogeny, fingerprinting, and provenance. PhyloLM's main text and the method sections of LLM DNA and of the provenance tester of Nikolic, Baluta and Saxena were read for it; that tester's null applies to one suspected child's aggregate agreement with its candidate parents, not to individual features, contrary to a first summary prepared for this revision. The other entries from the first search rest on their abstracts and are capped accordingly. A second search the same day, with each reported work checked against its primary record, found four works the first had missed: Stemma (Zhang, Safronov and Martin), posted on 28 July 2026, before this paper's first deposit on 20 August 2026, which takes its name from stemmatics and applies a feature-level version of the significance condition to shared wrong answers for pairwise provenance, read beyond its abstract; and, at abstract level, Gallagher et al. and Bruckner, who also build trees or fingerprints from outputs, and TokenPrint, whose identical-data result bears on §4. §10 cites all four and §7.1 cites Stemma. Three further works, brought by the author the same day, were read before citing: Kohli's panel study, which compares same-family with cross-family error correlation directly (§4, §5); antidistillation fingerprinting (§10); and Cisco's provenance explorer (§10). §10's statement that two methods build trees from outputs was corrected; the scope statement of the narrowed absence claim now records the three search passes. Because Stemma finds lineage signal in wrong answers that an unrelated model does not usually give, §3, §4 and §13 now say that a shared error counts when its independent origin is improbable, and §13 records Stemma's pairwise evidence beside engineered markers and teacher tests without treating one background model's reproduction rate as an independent-origin probability. (3) Abstract: "the present norm and worsening" lost "and worsening", which implied the trend through time that §4 says is not measured; the experiment sentence now names the similarity-based baseline. The page's search description said that models wrong together are "usually" wrong the same way; the source's means are 60% and 42%, so it now says "often". (4) §5: "does not hold in practice on the measurements in §4" was stronger than Kim et al., whose cross-provider statement is an inference from a regression model, not an analysis restricted to cross-provider pairs; restated. (5) One review observed that calling reviewers from different vendors "decorrelated" presumes the vendor-diversity assumption that §4's measurements do not establish. The body now says what was done: §3 defines working separately or blind as a procedure, not a measured independence of errors; "independent" and "independently" are replaced where they described reviewers or reviews (§§5, 7.1, 8, 9, 12, 13 and two notes in the References); "correlated" is replaced where it asserted an unmeasured dependence (§§9, 11.6, 12); and §9's "decorrelated apparatus" is now the multi-vendor review apparatus. Terminology note for the entries below: where they call seats, gates, lineages or review "decorrelated", the word records the practice's intent at the time, not a measured property; the dated entries are left as written. (6) References: the note on Gu et al. said students were "trained on 100% teacher output"; that holds for the sampling-based students only, and it is corrected; the same note no longer says that some watermark families "cannot" be learned, since the tested budgets showed weak or slow learning, not impossibility. (7) Found by the seat review of this revision and corrected before publication: §1 said models wrong together "tend to be" wrong the same way, now "often"; §4 said that vendor diversity buys less decorrelation as capability rises, which does not follow even from its premise, since correlation can rise both within and across vendors while the difference between them grows; it now says only that, if capability increases error correlation across vendors as well as within them, a cross-vendor panel of more capable models has more correlated errors than comparable weaker panels, and it attributes the measurement to the one study that made it, which did not isolate cross-provider pairs; §5 now says the measurements do not establish the diversity assumption, rather than that they do not support it, because §4 records that the regression attributed some agreement to shared vendor; §6's "That fraction is small now and growing" is marked as the paper's claim; §7.1's "That measurement is not yet available" became "No study examined here supplies that measurement", and §13's "not yet measurable" and §4's "do not yet measure" were scoped the same way; §11.8 now states the corpus-poisoning risk conditionally, as §6 does; and "correlated" and "dependence" were replaced where they asserted a dependence that was not measured; and §4 and §5 no longer single out any study as coming close to measuring the loss of correct minority findings, saying instead that the aggregate comparisons are consistent with it but do not count what voting discarded. (8) Found after the ninth round, when an outside research record on model lineage was checked against its primary sources: ErrorTrace (Zang et al., NeurIPS 2025), which Stemma cites and which the second search had missed, discounts a model family's shared errors by the error rates of other families for black-box attribution. §10 now describes it; §7.1 lists it among the work that exists and no longer implies that scoring against a population is new; §13 records its evidence; and the scope statement of the narrowed absence claim records that each search pass missed a close work. The same check found that antidistillation fingerprinting, which §10 already cited, distils and detects on code in its second version, and §7.1's account of what has been tested on code now says so. (9) Added before publication at the author's decision: Competing interests now names the drafting lineage (Claude, Anthropic), which also prepared the review briefs and collations, and records Anthropic's public allegations that the developers of four reviewing lineages (Moonshot, DeepSeek, Alibaba and Zhipu) used Claude's outputs to improve their own models; §12 states what that could mean for the review, including the alleged relaying of customer requests to Claude. The paper had not named its drafting lineage before, although the Introduction's byline did. Seat review of this revision: four AI reviewers of four developers (Kimi of Moonshot, Codex of OpenAI, DeepSeek, and GLM of Z.ai), joined in the eleventh round by a fifth (Gemini, of Google), read the changes, with the sources the new text relies on, under briefs byte-identical across reviewers, in nine rounds on 24 Sep 2026 and a tenth on 25 Sep 2026: four on the changes and their corrections, each later round going to the reviewers whose rows it answered, a fifth on the additions from the second search, four more on the corrections those additions required, the seventh also covering three works the author brought the same day, and a tenth, to two of the reviewers, on the additions in (8), whose three findings were corrected and put back to the reviewers that raised them, who closed them; an eleventh, on (9), went only to reviewers whose developers the allegations do not accuse: Codex and a Gemini-lineage seat (Google), whose first two runs ended without output at a permission prompt of its command-line tool, so that it read only the corrected text; Codex's four findings were corrected and put back to it, Codex closed three, and the fourth, which the Gemini seat also raised, was then corrected in their wording and not put back; the corrections listed in (7), and the refinements to (1) to (6) and (8), came from those rounds. Each reviewer's findings, and whether and how each was closed, are recorded row by row in the collation filed with the returns; a finding counts as closed there only when the reviewer that raised it says so, and where a finding concerned what a source says, the source passage was put to the reviewer first. None was closed by override or by vote. Three runs failed and were repeated with the failure recorded: GLM returned an empty answer twice at its default output limit, and one Codex run could not open its files. In the eighth round one complete Kimi run was discarded by its driver for a prompt-format error and not kept; the run that was kept used that round's standard prompt. In the tenth round the drivers wrapped the identical briefs differently: a standard preamble on the limits of the reviewers' file access went to Kimi in both of its passes, twice in the second by a driver error, and to Codex only in its second. The ninth round's five minor corrections were applied without a further round, and the collation marks them as not put back to the reviewers that raised them. The briefs, the changes, all retained returns and the collation are filed in the practice's report store; the discarded complete Kimi return is not preserved.

Reading copy, 22 Sep 2026 (front-matter rearrangement, no section text changed). To let a first-time reader reach the Abstract sooner, the front matter was reordered and de-duplicated: the full deposit-and-version-chain paragraph was collapsed into a one-line citation pointing here (its chain and hashes are the per-version entries in this appendix and the v1.32 entry below); the Competing interests section was moved to immediately after the Abstract, with a one-line disclosure kept above the Abstract so the interest is still stated before the argument; the author block was set on one line. No argument, abstract, or competing-interest text changed; the deposited version of record (v1.31) is unaffected and retains the original order.

Draft v1.32, 22 Sep 2026. Housekeeping only. Records the v1.31 deposit and corrects the front-matter paragraph's own claim about deposit state, which had gone stale the same way v1.26 recorded for v1.25 and v1.30 for v1.29: it said the concept DOI resolved to v1.29, true when v1.31 was written and false the moment v1.31 itself was deposited. The concept DOI now resolves to v1.31, deposited 22 Sep 2026 at 15:37 UTC, version DOI 10.5281/zenodo.22900116. Verified before writing, not asserted: the Zenodo record was fetched over its API after publication and reported v1.31 at that DOI with file Agentic-Stemmatics-Soons-2026-v1.31.pdf, 530,994 bytes, MD5 2e53f99e48acdd0e5b627d606f1462c6; the local file hashes to that identical MD5, and the served file downloaded after publication is byte-identical to the local one, so its SHA-256, 2952ce248e8076f7d73bb6f039a3161e6ce996ea26c8507b8c126916342e8142, is citable rather than assumed. The deposited PDF is a headless-Chrome print of this page's HTML source at research-site commit 255a974 with the site navigation and masthead suppressed by one print-only rule; 80 pages; nothing else differs from the page. The Zenodo record description was rewritten at this deposit to state the stakes before the argument, matching the opening paragraph added to §1 at v1.31. No argument text changed.

Draft v1.31, 22 Sep 2026. Presentation revised for readability; no claim changed. This revision is the published v1.30 text in a more readable form: sentences unbundled into claim, explanation and qualification; technical terms introduced after the problem they name; subheadings added within sections; the eight correspondences of §3 stated as prose with their status labels rather than as a table; em-dashes replaced by commas, colons, semicolons or parentheses throughout the prose and in this appendix, while the §9 frozen commitments and the References keep their original punctuation. The readability edit was produced by a model of the OpenAI (Codex) lineage on 21 Sep 2026 against v1.29, and was then checked for meaning drift on 22 Sep 2026 by three decorrelated seats, each blind to the editor's own list of changes: Kimi K3 (Moonshot) returned about forty-two medium findings and a long tail of minor ones, GLM (Z.ai) thirteen, and Qwen (Alibaba) one, on identical bytes. Every finding ran in one direction: the edit had quietly weakened claims the source asserted (proof to evidence, is carried by to can reveal, necessity to possibility, and a signature sentence of §6 dropped). The disagreement between seats was resolved against the source, not by vote: one hundred and nine softenings were reinstated from the v1.29 wording, one proposed reinstatement was rejected because the source did not support it, and the four frozen commitments of §9, including the success formula, were confirmed character for character against the rendered deposited PDF. The body was then compared word for word with v1.30 and found identical apart from the readability changes themselves. The three seat returns, the adjudication and the restoration record are filed in the practice's report store, and this revision is logged at the closure log. Two facts belong on the face of this entry. The seats' spread of one, thirteen and forty-two findings on the same bytes is itself an instance of the agreement trap this paper describes; a majority reading would have called the edit faithful. And an extraction artifact, a phrase split across a line break, produced four false "not in source" readings during the check, each resolved by reading the raw lines: the §8 catalogue, reproduced on the review of the paper that catalogues it. No existing claim changed. Two additions of content: an opening paragraph in §1 stating the stakes of the problem, which makes no claim not already made in §§4 and 5; and a related-work subsection in §10, on interacting agents and shared memory, citing Wei et al. (arXiv:2601.12538v2), which records agent-to-agent transmission through a shared memory or a shared training run as a candidate channel that §6's taxonomy does not yet include; it is recorded as a candidate, not measured, and it changes no existing claim.

Draft v1.30, 16 Sep 2026. Housekeeping only. Records the v1.29 deposit and corrects the front-matter paragraph's own claim about deposit state, which had gone stale the same way v1.26 recorded for v1.25: it said the concept DOI resolved to v1.25, true when v1.29 was written and false the moment v1.29 itself was deposited. The concept DOI now resolves to v1.29, deposited 16 Sep 2026, version DOI 10.5281/zenodo.22790100. Verified before writing, not asserted: the Zenodo record was fetched over its API after publication and reported v1.29 at that DOI with file Agentic-Stemmatics-Soons-2026-v1.29.pdf, 689,782 bytes, MD5 27dc0447ac4335168f73b646462e869a; the local file hashes to that identical MD5, and the served file downloaded after publication is byte-identical to the local one, so its SHA-256 (6ff674793d39e292adcbc6cc313953b3bcf1930c8b46c707380ea8153f307623) is citable rather than assumed. The deposited PDF is a print of this page's HTML source at research-site commit 72b4d27 with the site navigation suppressed; nothing else differs from the page. Recorded at the closure log's seq 26. No argument text changed.

Draft v1.29, 15 Sep 2026. Two sentences tightened after a file-access seat review of the companion Introduction found the same overstatements here. §4: "most of the agreement attributed to capability rather than shared vendor" held for one leaderboard only; on the other the vendor and architecture terms exceed the capability terms, and on both the intercept dominates; now stated per leaderboard. §6 reference note: "5% of fine-tuning data marked" now names the mixture; the remaining 95% was the same teacher's unmarked output; a separate run with human-text filler detects at 10%. No conclusion changed.

Draft v1.28, 15 Sep 2026. Three corrections, no new claims. §4: the correlated-errors agreement rate was given as "~60% in leaderboard settings"; the source reports a mean of 60% on one leaderboard and 42% on the other, and its cross-provider claim is an inference from its regression model, not a cross-provider-only analysis; both now stated as such. §4, closing corollary: "buys less decorrelation every year" asserted a time trend from a cross-sectional study; restated as a capability finding with the trend removed. §6: the marker-inheritance result was attributed to Pan et al. 2025 and described as "defeatable by paraphrase"; the inheritance result is Sander et al. 2024 and Gu et al. 2024, and ordinary paraphrase under a secret key does not remove the marker: it is removed by rule recovery, dilution, or neutralisation at generation time. Evidence trail: CVRR-LIT-2026-09-15-003 and CORRECTIONS-OWED-2026-09-15. Argument text changed in three sentences; no conclusion changed.

Draft v1.27, 24 Aug 2026. One line: v1.26's deposit-identity correction was independently confirmed by two decorrelated seats; Kimi (Moonshot) and Grok (xAI), run blind and in parallel, each fetching the Zenodo record and hashing the local file itself rather than trusting the drafting assistant's report. Both returned CONFIRM-CLOSED on every checkable claim (record existence and metadata, concept-DOI resolution, local SHA-256, and the MD5 cross-check) with zero discrepancies; Grok additionally queried Zenodo's versions-list endpoint directly, ruling out a newer, silently-skipped deposit. Recorded at the closure log's seq 15 (Kimi) and seq 16 (Grok).

Draft v1.26, 24 Aug 2026. Housekeeping only. Corrects the front-matter paragraph's own claim about deposit state, which had gone stale: it said the concept DOI resolved to v1.15 (20 Aug 2026), true when that sentence was written and false the moment a new deposit landed. The concept DOI now resolves to v1.25, deposited 24 Aug 2026, version DOI 10.5281/zenodo.22077929. Verified before writing, not asserted: the Zenodo record page was fetched directly and reported v1.25 at that DOI with file Agentic-Stemmatics-Soons-2026-v1.25.pdf, MD5 eed6aa85f9c1e3a1576320cca9ac1af9; the local working copy of that same PDF hashes to that identical MD5, so the deposited bytes are confirmed identical to the local copy, and its SHA-256 (9d0e40c1ae8c68290aef3fec656d8cf7a5fc951cbdb76878680985f9dc2a3c60) is now citable rather than assumed. This closes the "optional... no longer a blocker" item several companion documents had carried against this paper's earlier deposit lag. No argument text changed. (This entry's own closing asterisk (absent in the source draft, since the entry was drafted but never published until now) is supplied here, declared rather than silently repaired, the same discipline v1.24's NEW-K1 applied to an equivalent unbalanced marker.)

*Draft v1.25, 24 Aug 2026. The decorrelated gate on the v1.24 batch (Kimi, Moonshot lineage, fresh on every item) returned MINOR-REVISION (no BLOCK-grade finding) and this revision applies its five conditions.

The gate's record, in brief: items 8 and 8B REVISE, GR-7 and NEW-1 APPROVE as landed. The batch's load-bearing constraint (that §1's sharpening stays emission-scoped and leaves the interpretability question open) was ruled intact. The seat independently recounted §13's inventory by reading each span with its lead-in, catching both em-dash-phrased tags and all three line-wrapping tags, and confirmed Eight as then stated. It attacked item 8B's three premise corrections with its own controlled searches and upheld all three. Its rulings on the two withheld edits: §10's half-sentence IN (the body had moved while the formula stood still; this paper's own recorded defect class); §11.3's back-reference OUT, for two reasons, one of which was a spec defect no prior seat had caught: the candidate document's anchor names §11.3, but the sentence it targets lives in §5.

The five conditions, applied here: (1) the mode-persistence clause in §4's mechanism paragraph now carries the [verify] tag the seat judged the paper's own fail-closed convention demands; the definitional half of the paragraph needs no anchor, but "the direction survives" is an empirical claim about deployed models; §13 moves eight → nine accordingly, the ninth added by the seat's ruling rather than the author's drafting. (2) The under-trained-token tag's scope is widened to name model-discriminating power, the adjective that does the work in §7.1's edge argument and exceeded the declared abstract-level check. (3) §10's compressed formula now carries the census qualification. (4) The gate's NEW-K1 (an undeclared benign repair to a prior entry's markup) is declared in the v1.24 entry below. (5) The Land & Bartolo reference stands as its own entry. The gate's NEW-K3 (that the branch and commits the v1.24 entry cites were unverifiable from the served tree) is answered by the author: both commits confirmed present on the named branch. Its NEW-K4 (one sentence in §4's third-genesis paragraph entered from the discussion prose rather than either candidate spec) is accepted as content and recorded as provenance, here.

Confirmation of the five conditions by the raising seat is pending at this entry's writing; adoption waits on it.*

*Draft v1.24, 24 Aug 2026. Four items, landed as one batch; decorrelated gate (Kimi) running at this entry's writing: its record and any resulting rows append here, and adoption waits on it.

Items 8 and 8B (candidate documents on claude/agentic-stemmatics-ai-models-uitgwr, commits 30f70da and 8fa0d66): §4 gains the mechanism paragraph and the third genesis (shared tooling) beside attractor and inheritance; the "delve" voiding is renamed as the mechanism-import it was; §6's lectio facilior sentence is grounded in the training objective; §1's same-mechanism sentence gains its emission-side ground while leaving the interpretability question exactly as open as before; §7.1 gains tokenizer characters as tooling-descent validation edges, drawn as a second edge type; §7.3's drift probe gains the inversion; it runs on the very property that disqualifies the witness as an oracle. Item 8B corrected three premises of the discussion that produced it (no §6 hedge existed; the fluency slogan is the practice's, not this paper's; a "no separate truth channel" claim would have contradicted §1's deliberately open question): each correction landing on stronger existing text than the discussion assumed. Two optional edits (item 8's §10 half-sentence; 8B's §11.3 back-reference) are deliberately not applied, left to the gating seat.

The stratigraphy bullet in §6 is qualified (a row raised by a Grok seat against the companion, which found the same sentence at equal strength here): the ecology deposits candidates; the Leitfehler reading waits on §7.1's deferred instrument.

The References' "Still unresolved" inventory gains the §7.1 two-lineage record (a row raised by a DeepSeek seat confirming v1.23: a cited-but-not-reproduced record was absent from the paper's own inventory of exactly that class), and the [unverifiable] marker species is defined there; permanent limitations listed beside pending ones.

§13's inventory moves seven → eight (one tag added by the batch, none removed), with the arithmetic stated. One touch outside the declared sites went undeclared, and the gate caught it (its NEW-K1): the batch also repaired an unbalanced bold marker that the v1.23 entry below had carried since it was written; opened , closed with a single . Benign, the author's, and declared only here, at the gate's direction.

*Draft v1.23, 23 Aug 2026. Two rows from a scoped decorrelated gate on v1.22's §7.1 correction, both applied.

The seat was given the passage alone and no access to the cited papers: deliberately, since any excerpt of those would have been chosen by the lineage whose work it was checking. It answered the confabulation control with NO KNOWLEDGE of either paper and declined to touch the factual leg at all, which is the correct behaviour and the reason its other findings carry weight.

Its stronger row was that §7.1's reason for omitting a marker on the internal record (that a tag which can never be discharged is a defect rather than a disclosure) was a non sequitur. A marker warns about provenance as well as promising resolution, and a permanently unverifiable record is the strongest case for one. The passage now carries an explicit [unverifiable] marker, distinguished from [verify] because it will never close and §13 counts pending verifications.

Its other row read the paragraph as leaving its central absence claim self-contradictory: exhaustive in the prose, sampled in a nearby tag. Those are two different absence claims, 1,777 characters apart with an intervening change of subject, so the contradiction is not in the paper. But the seat is a careful reader and it conflated them, which is evidence enough that the text invited it; the two are now explicitly distinguished. Recorded rather than dismissed, with the cause named: the passage was served as a 2,704-byte excerpt, and that narrow framing made the two claims look adjacent. An over-narrow bundle manufactured the finding, exactly as an omission would have: the same failure in the opposite direction.*

*Draft v1.22, 23 Aug 2026. Two corrections, both to claims this paper made about its own evidence, and both surfaced by discharging [verify] tags rather than by a reader complaining.

§7.1's code-domain claim was FALSIFIED. The paragraph asserted, from an abstract-level check, that neither instrument citation reported a code domain, and concluded that code-specific application was an uncovered transfer. Full-text reading by two independent lineages found that Rawat et al. Appendix B.2 probes with MBPP code-generation prompts against mathematics-side distillation, and reports the method remains accurate under that mismatch. The Sun et al. half held: natural-language output only, confirmed including appendices. The correction is stated in the text rather than applied silently: an absence claim asserted from abstracts survived a revision before anyone read the appendices, and on that one point the paper had claimed the opposite of what was measured.

§9's public-artifacts claim was WITHDRAWN. The section said case reports are public documents where disclosure has completed, with a [verify] owing a pointer to them. Checked 23 Aug 2026: no case report and no reproduction artifact is published anywhere a stranger can fetch and run. The tag could not be discharged by finding out, only by publishing or by saying so; this revision says so. Every claim in §9 about what the practice produces now rests, visibly, on records the reader cannot open.

Both corrections concern the failure this paper theorises: a terminal claim, made once, carried forward, and false. Neither was caught by re-reading the draft.*

*Draft v1.21, 23 Aug 2026. Eight changes, landed as one atomic diff after a seven-round decorrelated review. §1 gains a constructive statement of what the program is for, bounded in the same paragraph and blocked from self-instantiation (§12, §13). §6 gains the objection-and-reply on capability improvement (that benchmarks measure central tendency while collation consumes dispersion, and that the one measured relationship moves the uncomfortable way) with the attractor channel and the recursion channel held apart and the degradation clause's remedies mapped to their channels. §7.1 records that the attribution instruments were measured on model output text, with no code domain reported at abstract level, so code application is an uncovered transfer. §9 gains the practice's other kind of evidence (a case record with per-finding evidence rungs) while refusing attribution, rate, and comparative value, and conditioning reader-checkability on a pending verification. §11.12 extends the denominator argument from the defect catalogue to this paper's own revision record. §11 gains limit 15, on the calibration of the framework's own skepticism: the commercial interest is declared four times over, but none of those declarations isolates its consequence for proportionality. §13 registers the §1 claim as a possibility claim and absorbs two new [verify] tags, taking the named inventory from seven to nine.

Review provenance, stated because this revision's closure record is unusually complete: nineteen rows were raised across five decorrelated lineages (Codex, Kimi, Grok, GLM, DeepSeek) and all nineteen were closed by the seats that raised them. None was closed by vote, override, or substitution, and no text in the diff shipped unseen by a decorrelated seat. Three of the fixes made during that process introduced fresh defects that later seats caught, and one insertion required five revisions before a seat would close it; both facts are recorded here rather than in the smoother form. A sixth commissioned round (Qwen) did not issue (it failed its corpus-receipt gate and reviewed the wrong object) and is excluded from that count.

This revision also logs a defect of its own, caught after the first build and before merge: the front-matter version paragraph above was left reading "Draft v1.20" while this revision was already v1.21, with its change enumeration stopping at v1.18; the third occurrence of the same stale-front-matter defect in this file, after the v1.16 lag two seats caught independently at v1.18 and the correction v1.19 exists to record. It was found by a peer reading the filing's own version label against its index row. The paragraph now enumerates every revision since the deposit.

This revision also logs a count error in the v1.20 note below: it referred to "§11's fifteen limits" when §11 carried fourteen. With §11.15 added here the count is fifteen: true only as of now, and recorded rather than silently trued. This is a further instance of this paper miscounting itself, in a second count distinct from the [verify] inventory whose six-revision history §13 documents.*

Draft v1.20, 22 Aug 2026. §13 only: the claims statement now carries the character-class scope the body already conceded three times (§4's per-item split, §6's "unresolved on the current evidence", §7.1's "which it is not yet"). One sentence added after the We-claim list ("The rule is presently earned for engineered markers and reference-based teacher tests only; per-item improbability for natural habits is not yet measurable, so natural-character stemmatics is program, not result (§6, §7.1)") and "executable today as planted-mark tracing" attached to the validation claim. Provenance of the change: a Qwen3.8 blind read of v1.18 returned this as its single strongest objection (verdict RESTRUCTURE; seat recorded as partially correlated (see the closure log, seq 6–7); a GLM-5.3 seat, given a four-passage excerpt with the objection to adjudicate, returned ANCHORED / transmission-failure / ONE SENTENCE and drafted the sentence adopted here verbatim) so the fix is outside-drafted, not author-drafted. Two source checks sharpened its verdict: §11's fifteen limits do NOT carry this scope (so nothing else transmitted it), and Qwen's premises are in §4 as claimed. GLM's residual (rebinding "progressively exact rather than presently identical" to the measurable channel) was deliberately NOT applied here; it remains an open editorial option. Restructure was declined on GLM's own test: body and claims do not contradict, and §13's claims are explicitly conditional.

Draft v1.19, 22 Aug 2026. Front matter only; no section text changed. The pre-Abstract version line had read "Draft v1.16, 20 Aug 2026" through two subsequent revisions: a stale version on the paper's own face, found independently by two review lineages on the same day (a Grok seat reviewing the companion Prolegomena, and a Codex seat reviewing this paper's library filing). The same stale line also claimed v1.15 is "textually identical to this one," which stopped being true at v1.16; the rewritten line now enumerates what changed since the deposit. One process note belongs in the record rather than out of it: the corrected paragraph was first applied in place to the v1.18 file, which briefly left two different texts carrying the v1.18 label while that label is hash-pinned in filed review records. This revision exists to end that: the fix carries a new version string, and the v1.18 file was restored to the exact bytes those records pin (SHA-256 6e848d7194195f9a0e716fc22c1a5f939268eaecfe307cdd9afbaef2c4af866a).

Draft v1.18, 22 Aug 2026. Cuts the §10 paragraph added in v1.17 and keeps one line from it, refolded into §3 as vocabulary: a generator expands the space of answers; it does not decide which answer matters. The line states the boundary set by §3's first row, so it now sits there rather than being introduced in §10 by way of an outside argument. What is dropped is the §10 framing around it: the observation that a humanities-for-AI argument is being made independently elsewhere, and an explicit disclaimer that this paper takes no position on machine cognition. Both were removed on the same reasoning: the method is indifferent to whether a witness understands anything, in the way stemmatics is indifferent to whether a scribe understood his exemplar; and is in fact better served when he did not, since a scribe who understands corrects silently and destroys the evidence. A disclaimer stating no position on the syntactic/semantic question still enters that debate in a minor key, and flags as live a question the testimony frame (§1) never asks. Reviewer note: a two-lineage panel on the v1.17 paragraph split; one seat advised editing it, one advised cutting it; cutting was chosen.

Draft v1.17, 21 Aug 2026. Added one paragraph to §10 recording that the humanities-for-AI argument is now made independently in popular and industry venues, noted as corroboration of reach and explicitly not as evidence (per §11.13), and folded in the compression above. Prompted by an IBM Technology video making the syntactic-vs-semantic argument; the distinction itself was deliberately not adopted. Superseded by v1.18, which cut the paragraph and kept the line.

Draft v1.16, 20 Aug 2026. Records the Zenodo deposit: concept DOI 10.5281/zenodo.22030516, version DOI 10.5281/zenodo.22030517 for the v1.15 file, CC BY 4.0, published 20 Aug 2026. The deposited PDF is Agentic-Stemmatics-Soons-2026-v1.15.pdf, 678,933 bytes, SHA-256 d70778bad4a6a6ed873f4db148f65bd053fa65b976e17217685a4f262dafcbd8. Zenodo freezes files on publication, so that byte sequence is now fixed and independently checkable against the record. The paper was not posted to arXiv: arXiv requires an endorsement from an existing author in the subject class for submitters without an institutional affiliation, which this practice does not have. That is a gate on visibility, not on priority, and it is recorded here rather than left as an unexplained absence.

Draft v1.15, 20 Aug 2026, NOT SUBMITTED. Adds the author's ORCID (0009-0000-3088-6373) to the byline, so the identifier travels with the PDF rather than living only in a hosting platform's metadata. This revision's note was written directly into this appendix rather than at the top of the paper: the practice v1.14 adopted after the changelog had twice re-accumulated in front of the Abstract.

Draft v1.14, 20 Aug 2026, NOT SUBMITTED. This revision moves the notes for v1.11, v1.12 and v1.13 into Appendix A, where the rest of the revision history already lives. They had re-accumulated at the top over three revisions, putting a reader's first page of this paper on its changelog rather than its argument: the same drift v1.11 corrected once already. Nothing in those notes was reworded or cut; only their position changed. Appendix A now carries the complete provenance record from v0.1 through v1.13.

Draft v1.13, 20 Aug 2026, NOT SUBMITTED. The closure log is now public (github.com/clvstra/agentic-stemmatics-closure-log), with force-pushes and deletions barred for administrators including the author (verified by attempting a rewrite and having it rejected, rather than by trusting the configuration) and with a standalone verifier anyone can run. This addresses the external-custody finding open since v0.4 without closing it: the repository is the author's own account, the second holder is the practice's other principal, and GitHub is a trusted party, so none of the three is independent. Per this paper's own rule the row closes only by the seat that raised it, so it is recorded as addressed and left open. §12 states the gain precisely and refuses the larger claim.

Draft v1.12, 20 Aug 2026, NOT SUBMITTED. The closure log's first entry was anchored to the Bitcoin blockchain this revision (txid 54274813...0145e0, block 963282, 20 Aug 2026 09:23:01 UTC), and §12 now records it together with an explicit account of what it does not buy. It does not close the standing external-custody finding, which has been open since v0.4 and is named in §12 and the Notes as still open: a hash proves what a record was, but only to a reader who already holds the record, so no external party has custody of anything. The log had also fallen out of use since v1.1; the two Grok closure rounds reported above were appended this revision, but retroactively: reconstructed from filed artifacts, marked as such in the log, and not anchored. A retroactively written entry is weaker evidence than a contemporaneous one, and the log now says so on its own face rather than presenting four entries as though they were all captured the same way.

Draft v1.11, 19 Aug 2026, NOT SUBMITTED. This revision moves the full revision history (previously the first thing a reader hit, before the Abstract) to Appendix A, after the References. Nothing in the history was reworded or cut in the move; only its position changed. See Appendix A for the complete provenance record from v0.1 through v1.10.

Draft v1.10, 19 Aug 2026, NOT SUBMITTED. UNGATED. Grok's closure round on v1.9 returned MINOR-REVISION (GROK-PAPER-V19-CLOSURE-2026-08-19.txt, filed by the operator from Grok's pasted return). All seven v1.7 leftovers closed except F-v03-6 (the closure-log external-custody anchor), which the round explicitly confirmed cannot be paid by wording and should stay open: no change made to it. Two new findings on material no prior round had reviewed as live text, both fixed here: (NF-v19-1) the v1.8 usus scribendi* [verify] tag was never folded into §13's named inventory, which kept saying six when the live body carried seven; a sixth consecutive wrong count, caught this time by a decorrelated seat rather than by either of this session's own full-text reads, now corrected to seven with the omission's cause stated; (NF-v19-2) the v1.8 header claimed to apply "the philology-literate human read... owed since v0.4," but the questions document that read was written for states plainly it needs a classicist and that no automated seat can supply it: and the source of the answers this session worked from was never confirmed before that claim was written. Asked directly this session; not yet answered. The v1.8 header is corrected in place (original text kept, correction stated above it) rather than silently rewritten, and the Notes item stays open until the source is confirmed. Grok also caught a citation hygiene error in the v1.9 header (this note, before this correction, cited a wrong date on its own source file): noted here as the kind of small self-account slip this paper keeps finding in itself.

Draft v1.9, 19 Aug 2026, NOT SUBMITTED. UNGATED. Grok's closure round on v1.7 returned MINOR-REVISION (GROK-PAPER-V17-CLOSURE-2026-08-19.txt), closing two Mediums it had left open across three prior rounds (F-v03-4, F-v03-5) and confirming NF-v06-1/2/5 fixed, but finding the v1.7 subtractive cut honestly stated and incompletely executed: three new Lows (NF-v17-1/2/3) where the header still narrated the withdrawn theory in the present tense, §8 pointed at a "§5 stopping line" the cut had removed, and the surviving §5 rule's trigger ("pooling correlated judgments") was narrower than the conditional it was meant to serve ("unmeasured or failing"), which would have let an unmeasured-but-uncorrelated pool through unchecked. All three are fixed here, along with three Lows left open across multiple prior rounds: F-v03-2 (the "§5 asymmetry" label mis-cited §5 and glued two separate dated instances into one (now separated, with the mis-cite removed), F-v03-7 (§9 overstated the PBM reviewing seats as having converged on a successor's design, when they converged only on discontinuing the original instrument) the successor's design is this practice's own), and NF-v06-4 (§10's "unobservable even to itself" closed a question §1 deliberately left open about the model's own internal state, not just the consumer's; a phrase that survived one prior round undetected because it wraps across a line break, invisible to a naive search, which Grok independently confirmed before fixing it). Not fixed, and not fixable by wording:* F-v03-6, the closure log's external-custody anchor; naming a party other than the author to hold or verify a copy is an operational decision this revision cannot make for the practice; it remains open in the Notes for the author, honestly, rather than closed by asserting an anchor that does not exist.

Draft v1.8, 19 Aug 2026, NOT SUBMITTED. UNGATED. Correction, added at v1.10: the paragraph below originally said this revision "applies the philology-literate human read... owed since v0.4." That overclaimed. PHILOLOGY-REVIEW-QUESTIONS-2026-08-18.md states plainly it is written for a classicist, textual critic, or codicologist, and that no automated seat can supply the read it asks for; the source of the answers this revision worked from was never confirmed before that sentence was written, and remains unconfirmed as of v1.10. The correct description is: this revision used the questions document as a checklist and applied the ten answers it was given against the live text, verifying what could be independently checked (bibliographic facts for two candidate citations) and leaving what could not be (whether those sources actually treat the term in question) as open [verify] tags. The owed classicist read itself (sourced, credentialed, attributable) remains open. What follows is the original v1.8 note, otherwise unchanged.

Draft v1.8, 19 Aug 2026, NOT SUBMITTED. UNGATED. This revision applies the philology-literate human read of §2–§4 that was owed since v0.4 (PHILOLOGY-REVIEW-QUESTIONS-2026-08-18.md), answered against the questions that document put to the reader. Of ten items plus a central question, six were already correctly implemented by v1.7 and needed no change (the central question; items 1, 5, 6, 8, 10): confirmed against the actual text, not assumed from the reviewer's own "Implemented" label. Four required real edits, applied here: (2) a disclaimer that Maas's own Leitfehler criterion was qualitative (unwahrscheinlich), not the P-notation this paper formalises it with; (3) a sentence locating the homoplasy/synapomorphy apparatus correctly (the logical distinction at the level classical stemmatics already required, the character-state machinery belonging only to the discipline's modern computational wing; (4) an explicit "structural analogy, not direct transposition" flag on the Bédier comparison, in both §2 and limit 5, since Bédier's finding was about stemma-counting specifically and the extension to panel independence is this paper's own; (9) the usus scribendi citation, which asserted Reynolds & Wilson as "the canonical account of the concept" without that specific claim ever having been confirmed) softened to cite Reynolds & Wilson for what is confirmed (a canonical general account of scribal transmission), with Pasquali's Storia della tradizione e critica del testo (Le Monnier, 1934) and Bischoff's Latin Palaeography (Cambridge UP, 1990) named as the reviewer's more precise candidate anchors for the term itself; both bibliographically verified this revision (author, title, publisher, date), neither yet confirmed to actually treat usus scribendi by name, so both carry [verify] rather than being cited as settled. Item 7 (a mechanical sweep of §§2–4 and §6 for "reconstruction" language doing work only emendatio* could license) was performed this revision: every instance was checked against §7.1's and §11.7's existing recensio/emendatio distinction, and none overreaches it; the sweep is closed clean, with no textual change required.

Draft v1.7, 19 Aug 2026, NOT SUBMITTED. UNGATED. This revision subtracts rather than adds. Across v1.3, v1.4 and v1.6 this paper attempted a general criterion separating decisions bound by §5's independence condition from operational choices exempt from it; each attempt was repaired by external review, and the third still had a hole. One reviewing seat noted that if the boundary needed a fourth structural iteration, the need itself was the finding. It did. §5's general theory and §9's application of it are cut (roughly eight thousand characters removed, the first net reduction in this document's history) and replaced by a stated withdrawal plus the one operational rule that survived: where a decision's justification depends on a factual premise settled by pooling correlated judgments, the condition binds that premise whatever the decision is called; where no such premise is load-bearing, it does not. §9 keeps the admission (the vote settled nothing about the clause; the discontinuation rests on the rule having been contested, an event rather than a finding) and drops the theory that was defending it. Nothing in §4, §6, or §7 (the paper's actual contributions) is touched. The prompt for this was a question about whether the review rounds had begun circling: they had, in one section, and the honest answer was to stop building there rather than to iterate a fourth time. v1.4 went to two seats independently. DeepSeek closed row I (open since v0.7 and the one this paper repeatedly called the one that mattered) judging the recursive decomposition to have fixed the displacement defect, and closed rows E and F as well. GLM verified the §5 literature live, the one check no text-only seat could run, and returned the sharpest finding of the round: the boundary was resting on the wrong citation. Judgment aggregation, in the formal literature, is defined by the objects aggregated (sets of yes/no judgments over connected propositions) and not by truth-tracking; its impossibility results concern consistency, not voter independence. Our appeal to it named a different axis than the one we needed. Confirmed independently against the primary source before acting, along with one supporting detail of GLM's own that did not hold: its claim that the entry's examples include normative propositions; they are factual and procedural. §5 is rewritten accordingly: the boundary is the older descriptive/prescriptive one, the independence requirement rides on Condorcet's Jury Theorem and the dependent-vote results (with two precisions GLM's check forced; sufficiently low correlation* rather than strict independence, and binary votes rather than pooled estimates), and the judgment-aggregation appeal is withdrawn as a misuse rather than quietly dropped. §5 also now concedes what it is not inventing: Popper demarcates theories rather than sorting judgments, the descriptive/prescriptive line is Hume's, and the recursion is Quine–Duhem's structure with the stopping line in the classical observation-statement role.

Both seats independently reached the same two fixes, which is convergence worth recording: the stopping line must test a premise's content rather than its mode of access (a record of a pooled judgment makes the pooling observable, never the conclusion observational), and §9's claim that correlation strengthened the non-convergence premise had to go. That claim was ours and we liked it; both seats independently identified it as having the every-outcome-confirms shape §11.13 already catalogues (correlated agreement discounted, correlated disagreement upgraded, no model making either direction wrong) and GLM added the arithmetic, that at three seats and one draw each a 2-to-1 split occurs around 38% of the time even on a strongly-peaked reading. Withdrawn to the narrow claim correlation does not weaken the observation. Row B is fixed by broadening §3's oracle definition to include normative-text consultation (both seats chose this over narrowing §8), with the three carry-throughs GLM identified and we had missed: §7.3's ratchet, §11.2's enumeration, and §9's "grading oracle," which was a third sense of the term hiding one section away. §9's own premises are reclassified in the same pass. Rows B (pending confirmation), D, G, J remain; J is now honestly bounded rather than argued, and §9's self-application is demoted from demonstration to illustration because the disputed clause cannot be reproduced. v1.3 went to DeepSeek, which returned MAJOR CONCERNS (narrower than v1.1) and found a real hole in the boundary v1.3 had just inserted to fix an earlier hole. The §5 test as first written: does the judgment forbid some state of the evidence, or only prescribe an action? : scrutinised the speech act and not the premises underneath it, so a governance decision resting on a pooled, correlated factual premise would pass at the top while laundering a truth-claim one level down. That is the same defect the boundary was written to close, displaced rather than removed, and the criterion did not survive contact. §5 then stated the test recursively, with a counterfactual criterion for which premises are load-bearing and an explicit stopping point so it would not collapse into "everything is truth-reconstruction": this recursive apparatus is exactly what v1.7 later withdrew (see the note above); it is described here in the past tense because it no longer exists in the live text. §9 then applied the decomposition to its own discontinuation and, at the time, was judged to pass it on the merits: notably because its second premise used the panel's disagreement as directly observable data rather than any seat's conclusion as authority, a "correlation strengthens rather than undermines" framing that DeepSeek and GLM later independently flagged as an every-outcome-confirms claim (see below) and that does not appear in live §9. DeepSeek separately confirmed, by its own counterfactual test, that §9's cost/risk justification is genuinely independent of the withdrawn clause-interpretation claim. Two of its other rows are also addressed: the disputed clause is still not quotable (internal record, cited not reproduced) but §9 now says so plainly and points out that the ambiguity is visible in the paraphrase it does give, and §10's unscoped interpretability absence-claim is scoped per §8's own rule. Rows B, D, F, G remain open and are not text-fixable. A third independent count of the [verify] tags landed on six, matching. A consolidated review report (source unattributed at drafting, later confirmed as the operator's own synthesis of a philology questions brief, a governance briefing, and a Qwen-authored memorandum whose bibliography was found substantially degraded and is not itself cited here) proposed sixteen findings; ten are applied this revision after independent checking, not on the report's word. Verified directly: Timpanaro's actual argument (Bentley, Wolf, pre-Lachmannian origins (confirmed against the book's own chapter contents, not just its existence); the Bédier/Cerquiglini conflation (now separated as opposed rather than continuous reactions); usus scribendi's canonical source (Reynolds & Wilson, replacing a blog citation) chapter locator not claimed, since that detail could not be confirmed). Applied on the strength of their own reasoning: the codex descriptus table row renamed rather than stretched past its logical fit; the hyparchetype disclaimer moved to first use; a closing methodological statement on when the philological frame is dropped, not just worn. The consequential fix at the time was §5/ §9: a boundary (judgment aggregation requiring independence versus preference/risk aggregation that does not, then grounded in the formal judgment-aggregation literature) replacing the unfalsifiable "governance decision" label DeepSeek's row I correctly attacked, with §9 no longer defending its 2-to-1 majority by relabelling it. That specific grounding was itself later found to be a misuse of the cited literature (GLM; see the v1.5 → v1.6 note below) and the whole boundary apparatus was withdrawn at v1.7 (see the note above). What survives into the live text is narrower: §9 withdraws the claim that the vote settled the clause's meaning, leaves that interpretation an open residual, and rests the discontinuation on independent cost/risk grounds that never depended on the vote being right.

Provenance of v1.1 → v1.2 (18 Aug 2026): §12 rewritten that revision to state plainly that its own v0.3 closure-log commitment went unbuilt across seven revisions, that the historical record will never be covered retroactively, and that a minimal version now exists (hash-chained, started at v1.1, no external custody yet) logging the invocation of this very DeepSeek round as its first entry. v1.0 corrected the stale Notes-for-author section that a second unattributed review (no lineage, no envelope, no terminal marker; still unidentified when asked a second time) had read as a live "hard block": the COI/affiliation metadata itself had been correct since v0.7. That same review returned eleven further findings, ten applied. Two seated reviews of v1.0 then followed: Qwen (formalising its own v0.4 rows and v0.7 assessment; verdict CLOSED on both) and GLM (its first full-text pass after two v0.4 verdicts it had itself disclosed as bundle artifacts, withdrawn this round; verdict MINOR-REVISION). GLM's live primary-source verification caught a real attribution error (the watermark-inheritance citation's first author is Pan, not Zhao, independently confirmed against the ACL Anthology entry before correcting it) and a fifth consecutive wrong [verify]-tag count, this time for a third, distinct reason: the tag lives inside a sentence that wraps across a line break, invisible to every line-based automated search this session ran, including the one that produced the immediately preceding wrong count. Two independent full-text reads, not automated recounts, caught it; both had also independently found the correct number where four prior grep-based passes had not. Also applied: GLM's finding that the restructured §5 exemplar list named three systems but left the discard-singletons mechanism with no cross-vendor exemplar, stated as a gap rather than smoothed over; a softened §8 absence-claim commitment, since four instances did not carry the scope §8 promised of them; and, on GLM's own suggestion, a correlation-weighting sentence ported from §12 into §9, pricing the PBM discontinuation majority at the same epistemic grade §12 already applies to this paper's own four-seat convergence; this does not resolve DeepSeek's row I, it prices it honestly rather than leaving it merely disclosed. Earlier fixes carried forward: the §4/§5 minority-filtering contradiction; the §6 watermark-versus-natural-habit correction; the §7.1 scoping sentence; the §10 artifact-claim correction; undefined DeepSeek row letters and "no response envelope" glossed on use; and a Cursor reference repaired after an earlier cleanup script had broken it, with an unaudited figure dropped rather than left as an unused floating claim. Not applied: the review's own commentary and philology notes, since its lineage is still unconfirmed. DeepSeek's v0.7 findings (labelled B, E, F, G, I, J in that round's own numbering, not otherwise defined in this text: an oracle-citation framing question, an overstrong interpretability claim, the unauthenticated review history, the append-only log's completeness gap, the majority-configuration tension in §9, and the unquoted governing clause) remain untouched; all six want a record or measurement outside the author's sole control that redrafting text cannot supply. Provenance of v0.1: four separately-commissioned model seats of distinct vendor lineages ("distinct lineage" is a commissioning fact, not a measured independence claim (§11.6, §12)) (Grok, Codex: driven; GLM, Qwen: operator-carried, no response envelope, so no machine-readable record of which served model, token count, or run identity produced the return, only the operator's own filed header saying so) returned a unanimous MAJOR-REVISION; v0.2 was their disposition, machine-verified by nine checker agents against all 209 extracted rows.

Provenance of v0.2 → v0.3 (17 Aug 2026): four bounded additions from the PBM critique arc (§8, §10, §12 (new paragraphs) and §9 (status note)) following verification, a three-lineage external seat round, and an informed audit. Full record: PBM-DISCONTINUATION-RECORD-2026-08-17.md, PAPER-V03-STAGED-EDITS-PBM-ARC-2026-08-17.md.

Provenance of v0.3 → v0.4 (same day): v0.3 went through its first closure round against six seat-passes; Grok, Codex, DeepSeek (full/delta access) and GLM twice, once driven on a trimmed bundle and once hand-carried with live web access, plus a full-text hand-carried Qwen pass. Verdicts split cleanly by access: every seat given the full disposition text landed MINOR-REVISION or better on the pre-existing argument (Grok, Qwen; Codex closed 13 of 22 old rows outright); GLM's two MAJOR-REVISION verdicts are substantially an artifact of a bundle design that could not supply disposition text, disclosed as such by GLM itself both times (structural, not content). Six concrete, verified fixes applied this revision: (1) §8's "internally-correlated"/"correlated internal passes" language dropped (asserting correlation without measuring it, even of a critique's own investigations, is the exact discipline this paper argues against; (2) §8's "all three lineages independently rejected" corrected to state the actual 2-firm/1-partial split (Qwen's own verdict on the underlying question was PARTIAL, not a rejection) a Grok fact-check against the filed seat record, not a stylistic softening); (3) §8's citations named directly (SARIF, sarif-spec#120, in-toto/attestation#77) rather than by class, and the four verification inaccuracies' provenance attributed to the verification pass itself; (4) §10's in-toto citation names the repository in reader-facing prose, not only in an internal [verify] tag that would not survive to publication; (5) §10's SARIF timeline precised (2019 Committee Specification, 2020 OASIS Standard) and its kind enum given as six values with three named tokens, resolving both the earlier trichotomy imprecision and a citation-repo ambiguity a hand-carried GLM pass caught by live primary-source verification; (6) §12's closure-log commitment gained a self-aware caveat; the anchor itself is unnamed and a further instance of the same problem if chosen unilaterally by the author. Two items explicitly NOT touched this revision, named rather than silently deferred: formal affiliation/correspondence/COI metadata (Qwen's new Finding 27 (needs real values, not placeholder text) and the remaining [verify] tags (a citation-research task at a different scale, now underway) see the v0.4 → v0.5 note immediately below for what has been resolved so far).

Provenance of v0.4 → v0.5 (18 Aug 2026): the v0.4 closure round returned six verdicts (Codex MAJOR-REVISION; Grok, DeepSeek, Kimi, GLM, Qwen MINOR-REVISION), converging on a short list of small, confirmed defects; a self-contradictory [verify]-tag count (header said 40, notes said 38; three independent counting methods, Grok/Kimi/GLM/Qwen, converged on 35, not yet reconciled into the text below pending the fuller citation pass this note describes), an Abstract sentence that kept the unmeasured-correlation language §8's fix had dropped elsewhere (fixed below), and a §9 sentence that said "was verified" where §8 says "substantially settled, with edge corrections... not settled" (fixed below, matched to §8's own qualifier). One finding was not mechanical: Kimi, reading the paper cold with no framing inherited from the other five seats, found that §4's corrected inference rule applied its own improbability-under-independent-genesis condition to the error channel but not to the idiosyncrasy channel it replaces errors with; the cited 97% attribution figure measures discriminability (can models be told apart), not improbability (would an independent model develop the same idiosyncrasy by chance), and the paper's own §6 discussion of RLHF-driven vocabulary convergence (e.g. "delve") is a real, citable candidate case of exactly the convergent-attractor problem this channel was supposed to avoid. §4 is revised below to apply the improbability condition per idiosyncrasy rather than to the channel wholesale, splitting engineered markers (watermarks: improbable by construction) from naturally-occurring idiosyncrasies (which need the same per-item scrutiny errors get, not yet uniformly available). A parallel six-thread citation-verification pass checked roughly twenty [verify]-tagged claims against primary sources; most held, several needed real correction (not just sourcing): including a watermark-inheritance citation that had conflated two unrelated papers, one of which argues the opposite of what it was cited for, and a wrong technical term ("dictionary ghost words," which properly denotes an accidental lexicographic error, not the deliberate copyright trap the paper meant; the correct term is "Mountweazel"). The full bibliography and remaining tag resolution are a separate, still-in-progress pass; this revision applies only the items resolved so far plus the three items above.

The research field this paper proposes has a reference site of its own (definition, open problems, and the map of adjacent literatures) at agenticstemmatics.org. This page is part of the research section of CounterProof Research, a practice of Clavestra Capital Limited.