CounterProof Research · Paper
CounterProof Research — preprint, not peer-reviewed. Published version of record: doi:10.5281/zenodo.22030516 (CC BY 4.0).

Agentic Stemmatics: Collation, Provenance, and Conservation for an Emerging Machine Textual Tradition

A field report and a research program.

L. J. Soons Chief Strategy Officer, CounterProof Research Ltd ORCID: 0009-0000-3088-6373 Correspondence: admin@counterproof.io DOI: 10.5281/zenodo.22030516

Competing interests. The author is an officer of CounterProof Research Ltd, a commercial practice that performs and sells adversarial review services using the method this paper describes. The practice is the source of every internal observation reported here, including the field record in §8, the commitment record in §9, and the review process in §12; it stands to benefit from the method's adoption. No external funding supported this work. The disclosure is placed before the argument rather than after it because the paper's own position is that an interest should be stated before it is argued around, and because several of the limitations in §11 — selection bias, publication bias, the unaudited status of the internal record, and the author's role as adjudicator of his own review panels — follow directly from this relationship rather than being incidental to it.

Draft v1.16 — 20 Aug 2026. Deposited as a preprint on Zenodo, CC BY 4.0. The DOI above is the concept DOI and always resolves to the most recent deposited version; the version deposited on 20 Aug 2026 is v1.15, which is textually identical to this one except that it does not carry the DOI (it could not — the identifier did not exist until the deposit completed). Revision history is in Appendix A.

External claims carry [verify]; internal anecdotes are labelled illustrations, not evidence. The author operates a commercial practice built on the method described — stated first because the paper argues that interests must be disclosed before they are argued around.


Abstract

Large language models emit testimony: fluent claims whose warrant lies outside the act of generation. We propose that the right science for such output is the one philology built for corrupt textual transmission — stemmatics — but not by the identity claim an earlier draft of this paper made and four separately-commissioned reviewers correctly returned for major revision. The defensible thesis is time-indexed and program-shaped: (1) the analogy between model lineages and manuscript witnesses was premature as a description of the 2023–24 model ecology, where shared model errors are — on the measured evidence presently available — dominated by convergent attractors rather than inheritance; (2) it is becoming exact, because model output now re-enters training corpora through measurable transmission channels — direct distillation (with reported watermark-inheritance evidence), ambient corpus contamination, deliberate synthetic data, and human-mediated diffusion — so that the ecology is growing a genuine descent tradition; and (3) the genealogical signal in models is carried not by shared errors (philology's classical conjunctive errors) but by shared idiosyncrasies, for which the attribution and distillation-detection literatures report candidate instruments. From this corrected foundation we derive: a conditional aggregation result — majority voting over model reviewers is unsafe specifically where independence is unmeasured or failing, a condition the correlated-errors literature indicates is the present norm and worsening; an executable experiment — reconstructing the stemma of the model ecosystem from output idiosyncrasies and validating it against known model phylogenies; and a conservation program (archived model vintages as sealed witnesses, provenance strata, and a degradation clause for multi-model review panels). We report the failure record of our own instruments, including the reviewed failure of this paper's previous draft, whose disposition table is part of the method, and a second, dated instance from a sibling instrument of this practice — external verification settling a claim that internal agreement alone could not, and its own governance clause exited visibly rather than laundered. Commitment values for the practice's first controlled comparison were frozen and quoted from their commitment record; that instrument has since been discontinued on critique, before any run — a decision we report rather than paper over. We claim a framing, a corrected inference rule, and a program — not demonstrated efficacy.


1. Introduction: the output is a witness — under conditions

A conventional program can, in principle, be read; its behaviour derives from its text. A language model's output cannot be read this way. It is a claim — "this code is correct," "no such vulnerability exists" — and the claim's truth is not reliably recoverable from the fluency or internal consistency of the text asserting it. Structurally, the mechanism that emits a well-grounded claim and the mechanism that emits a baseless one are the same mechanism; whether any internal mark distinguishes them is an open question for interpretability, not a property available to the consumer of the output [Turpin, Michael, Perez & Bowman, "Language Models Don't Always Say What They Think," NeurIPS 2023, arXiv:2305.04388; Lanham et al., "Measuring Faithfulness in Chain-of-Thought Reasoning," Anthropic, 2023].

We therefore treat model outputs as witness testimony — the witness proper being the model lineage, sampled repeatedly, of which any single output is one draw (§3) — to be weighed, collated, and attributed rather than trusted or merely averaged. The discipline that industrialised witness-weighing for text is philology, and its genealogical wing — stemmatics — is the specific craft of recovering truth from multiple, individually corrupt, partially dependent witnesses.

What changed since the previous draft. Version 0.1 of this paper asserted the mapping was "structurally identical" and treated stemmatics' inference engine as directly transferable. Four separately-commissioned reviewers unanimously returned that draft for major revision, and their two deepest findings — that the paper never performed a distinctively stemmatic operation, and that shared model errors indicate convergence rather than descent — were correct. This version is built on those findings rather than around them. The result is a weaker claim about the present, a stronger claim about the trajectory, a corrected inference rule, and an experiment with available ground truth.

Contributions. (i) A conditioned mapping between stemmatics and multi-model verification, with its failure points stated as precisely as its correspondences (§3). (ii) The homoplasy correction: why the classical conjunctive-error inference does not transfer to models, and the corrected rule — model ancestry rides shared idiosyncrasies that are individually improbable under independent genesis, not shared errors — with the measured evidence for the attractor half, and an explicit account of which idiosyncrasies do and do not presently clear that bar (§4). (iii) A conditional aggregation result for multi-model review, grounded in the external correlated-errors literature rather than in our internal anecdotes, with the failure modes of our own adjudicative alternative stated symmetrically (§5). (iv) The corpus-recursion argument: the measurable channels by which the model ecology is becoming a textual tradition, and what follows for dating, conservation, and security (§6). (v) An executable program: the stemma-of-models experiment against known lineages; an empirical independence metric for review panels; and a conservation architecture with a degradation clause (§7). (vi) A field record of instrument failures, including this paper's own reviewed failure and a dated external instance from a sibling artifact, offered as evidence for the framing and explicitly not for efficacy (§8, §12). (vii) The frozen commitment record of the practice's first controlled comparison, quoted with its own must-not-claim boundaries, and its subsequent discontinuation, reported with the same discipline (§9).

We do not claim the method produces better code. That claim is reserved for a registered experiment whose commitment values existed and whose run never occurred (§9).

2. Stemmatics, stated with its conditions attached

Classical texts survive as copies of copies; the original — the archetype — is lost. The genealogical method conventionally attributed to Karl Lachmann — an attribution Timpanaro qualifies not by disputing that the method exists under that name, but by demonstrating that Lachmann adapted and formalised existing classical philological practice (the humanist tradition of emendatio ope codicum running through Bentley, and F. A. Wolf's Prolegomena) rather than inventing the method ex nihilo [Timpanaro, The Genesis of Lachmann's Method, trans. Most, Univ. of Chicago Press, 2005; orig. La genesi del metodo del Lachmann, 1963 — the contestation is historiographical, about originality, not ontological, about whether the method is real] — reconstructs archetypal readings from the pattern of agreement and disagreement among witnesses. Its engine is the conjunctive error: a shared error so arbitrary that independent double-genesis is improbable, which therefore evidences common ancestry.

Two conditions, present in the discipline from the start, matter for everything that follows.

First, the improbability condition. Not every shared error is conjunctive. Philology explicitly excludes polygenetic errors — banalisations, easy slips, orthographic modernisation — precisely because independent scribes converge on them. Only significant errors (Maas's Leitfehler, coined by analogy to geology's index fossils and subdivided into separative and conjunctive errors) carry genealogical signal [Maas, Textkritik, 1927; English Textual Criticism, trans. Flower, Clarendon, 1958]. The inference was never "shared error ⇒ descent"; it was "shared feature ⇒ descent only where P(shared | independent genesis) is negligible" — Maas's own criterion was qualitative (unwahrscheinlich, improbable, under independent origin); the P-notation is this paper's own formalisation for computational operationalisation, not language Maas himself used. Biology faces the identical problem as homoplasy (convergence) versus synapomorphy (shared derived character). The distinction itself is imported at the logical level, where classical stemmatics already required it, holistically and editor-dependently, without any formal apparatus; the machinery that makes it computable — character-state matrices, cladistic parsimony — belongs to the modern computational wing of the discipline, and computational stemmatology explicitly shares phylogenetics' machinery for it, including benchmark evaluation against artificial ground-truth traditions [Roelli (ed.), Handbook of Stemmatology, De Gruyter, 2020; Roos & Heikkilä, "Evaluating methods for computer-assisted stemmatology using artificial benchmark data sets," 2009].

Second, the discipline's own deflationary history. Stemmatics is not an unchallenged triumph. Bédier observed that reconstructed trees came out suspiciously two-branched and concluded the method partly manufactures the structure it claims to find [Bédier, 1928 — of 110 published stemmata surveyed, 105 were two-branched]. The field's response split into two directions the paper's earlier drafts ran together as one trajectory, and should not: Bédier himself retreated to best-manuscript editing — he still sought a single recoverable text, abandoning only the genealogical method as the instrument for reaching it — while the later New Philology [Cerquiglini, In Praise of the Variant, trans. Wing, Johns Hopkins, 1999, orig. Eloge de la variante, 1989] took the opposite turn, abandoning the goal of a single archetype entirely in favour of variance as the object of study. We cite both as the discipline's own deflationary history, not as one continuous move. We use stemmatics as an engine with its internal critiques attached. The Bédier problem reappears in our setting by structural analogy, not direct transposition — both cases involve a method manufacturing confidence in the very structure it cannot verify from inside — as the impossibility of fully verifying panel independence from inside the panel (§11).

Four further terms recur [definitions against standard handbooks — verify]: codex descriptus (a copy of an extant witness, eliminated as informationally redundant); contaminatio (a witness copied from multiple exemplars — mixed ancestry, the hardest corruption to detect); lectio difficilior potior (the harder reading is likelier original, because copying drifts toward the easy); and the hyparchetype (a shared corrupt intermediate ancestor, behind which no collation can see).

A note on the frame's status. We do not claim a 1:1 ontological identity between manuscripts and LLMs. We adopt the inferential logic of stemmatics -- its handling of transmission, corruption, independence, and the improbability condition -- as a formal epistemological framework for machine testimony. Where the classical vocabulary maps precisely onto the computational reality (contamination, stratigraphy, the improbability condition), we use it. Where it does not (strict eliminatio, emendatio), we discard the Latin and keep only the logical operation -- §3's table is revised in several rows on exactly this basis. The value of the philological frame is not its vocabulary; it is a centuries-tested discipline for managing correlated witnesses, and a paper using that discipline should be willing to drop the costume the moment the costume stops fitting.

3. The mapping, conditioned

The table below replaces the previous draft's. Each row now carries its status; three rows that reviewers showed to be false or forced have been corrected or removed, and the corrections are substantive, not cosmetic.

Stemmatics Multi-model verification Status of the row
Lost archetype Ground truth about the artifact Conditioned. Holds only for decidable questions — throughout this paper that names the operational sense, a question some oracle settles — an execution, a type-check, a test, a proof, or consultation of a normative text where the question is what that text says (a specification's enum, an issue's status), which settles the content of the text and never whether the text is correct or whether some claim about the world holds — and not the computability-theoretic sense of the term; the general property is undecidable there, and no argument here needs it to be otherwise. For judgment questions (severity, design soundness) there may be no single archetype — reviewers may answer different legitimate questions, not corrupt one truth. And unlike the philologist's, our archetype is often extant but expensive: the code can be read, tests can be run. Where an oracle exists, collation is subordinate to it. This bounds the framework's domain (§11).
Manuscript witness A model lineage, sampled repeatedly Corrected. The previous draft mapped the witness to a single output, which its own limits section contradicted: a model returns different verdicts on identical input. The atomic unit is the lineage-as-distribution; a single output is one draw from a witness, not a witness. Consequences for blocking rules are drawn in §5 (a block is one draw, to be established, not a verdict) and §11.3.
Conjunctive error ⇒ shared ancestry ~~Shared model error ⇒ shared training lineage~~ Replaced. The classical engine does not transfer on the error channel — §4. The corrected rule: significant shared idiosyncrasies ⇒ lineage.
Witness redundancy / sampling variance Second same-family output: discounted for measured correlation, not eliminated Renamed. Philology's strict eliminatio codicum descriptorum removes a copy because it adds zero genealogical information. A second LLM sample is not that: it adds distributional information about the lineage's variance, so the operation is discounting for measured correlation, not elimination. The logical character of the operation has changed enough that we drop the Latin term for this row rather than stretch it to cover a different operation. What eliminatio correctly warns against — counting correlated witnesses as independent — is a weighting error, and that warning we keep.
Contaminatio Cross-seat contamination: feeding one reviewer another's findings Holds, narrowly. Illustrated (not evidenced) in our field record (§8); review survival is not offered as warrant (§13). A shared evidence bundle is a milder relative — common inputs rather than mixed exemplars — and the two should not be lumped.
Independent agreement ⇒ likely truth Decorrelated agreement ⇒ likely truth Conditioned on measured independence — which cannot be assumed from vendor identity (§4, §5) and cannot be fully verified from inside the panel (§11).
Lectio difficilior ~~Distrust the fluent output~~ Removed as a synchronic rule; re-derived diachronically in §6. As a rule for choosing between two present outputs it is decorative — "difficulty" is not a defined signal, and scribal error is not uniformly directional [verify]. But as a claim about iterated transmission it becomes exact: recursive training on model output erodes distributional tails (§6; Shumailov et al., Nature 631:755, 2024) as scribal banalisation erodes hard readings. The row earns its life at the level of the tradition, not the sample.
Hyparchetype Shared training corpus Functional analogue, not a strict mapping. Classically a hyparchetype is a specific, lost, reconstructible intermediate manuscript; we extend the term to a shared training corpus, which is not a single ancestor. The extension is deliberate: we use "hyparchetype" for the risk class — correlated blind spots behind which collation cannot see — because that risk class is what matters here, not the entity-status of the common source. We acknowledge this goes beyond the term's strict codicological meaning, stated here rather than left implicit (§6, §11).

Two admissions govern the table's use. The generic force of several rows is common-mode-failure analysis, which reliability engineering possesses without Latin; what the philological frame adds beyond vocabulary is (a) the diachronic apparatus — descent, strata, contamination, conservation — which becomes literal in §6, and (b) a worked historical case in which output-level science proved adequate for reconstruction-grade warrant without mechanism-level access — in that domain (§10; whether the adequacy transfers here is part of what §7 must show). And the previous draft's most damning review finding — that no distinctively stemmatic operation was ever performed — is accepted and answered not by rewording but by §7, where the operations are specified as experiments.

4. The homoplasy correction: where lineage signal actually lives

The deepest objection raised against the previous draft deserves full statement. Scribal conjunctive errors evidence descent because they are improbable coincidences under a copy process. Model errors are not like this: two models trained on overlapping data toward similar objectives will converge on the same plausible-but-wrong answers as attractors — the mode of a shared objective — with no copying involved. Shared falsehood therefore fails to evidence shared ancestry. As stated, this threatened the framework's engine.

The resolution has two measured halves.

The attractor half is real, and externally measured. The correlated-errors literature reports, across hundreds of models, that when two models both err they agree on the same wrong answer at rates far above chance (~60% in leaderboard settings); that error correlation persists across distinct architectures and providers; and — most consequentially — that more capable models exhibit more correlated errors [Kim, Garg, Peng & Garg, "Correlated Errors in Large Language Models," ICML 2025, PMLR 267:30038-30066, arXiv:2506.07962 - over 350 models across two leaderboards and a resume-screening task]. A capability-controlled audit finds majority vote beats the best single ensemble member in under 10% of canonical three-model subsets, an effect its author characterises as modest and configuration-dependent rather than decisive [Kim, "Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles," arXiv:2607.20768, 2026 - preprint, not peer-reviewed at time of writing; 31,900 subsets over 30 models on MMLU-Pro]. That consensus selection thereby filters out minority-correct answers is our inference from the association between shared error overlap and lost gain, not a finding either paper states. So on the error channel, the objection stands: shared errors are predominantly polygenetic. Philology's own exclusion rule, applied honestly, discards most of them as genealogically void.

But the idiosyncrasy channel carries real, measured signal — establishing which idiosyncrasies clear the same bar errors must clear took a second look, prompted by an external reviewer who found the first version of this section had not done it. Models exhibit stable, model-specific idiosyncrasies — word-level distributions, structural habits, formatting signatures — sufficient for 97.1% five-way attribution between major systems, robust to paraphrase [Sun, Yin, Xu, Kolter & Liu, "Idiosyncrasies in Large Language Models," ICML 2025, arXiv:2502.12150]. That figure is discriminability: each model's idiosyncrasy profile differs enough from its neighbours' to classify unseen text. Discriminability is not the improbability-under-independent-genesis condition §2 requires, and at least one prominent idiosyncrasy fails it on present evidence: lexical overrepresentation of words like "delve" is measured across unrelated post-2022 models and, independently, in human scientific writing, with a candidate common cause in shared RLHF methodology rather than shared lineage [Juzek & Ward, "Why Does ChatGPT 'Delve' So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models," COLING 2025, arXiv:2412.11385 — the paper leaves the causal mechanism open, which is why we say candidate]. That is the same convergent-objective attractor §4's first half already uses to void shared errors as evidence, applied here to a habit instead of a mistake — and it means the idiosyncrasy channel inherits the error channel's own problem unless checked per item.

The corrected rule therefore needs the improbability condition applied idiosyncrasy-by-idiosyncrasy, not granted to the channel wholesale, and the channel splits into two sub-classes that earn it on different grounds. Engineered markers — a statistical mark planted in a teacher's outputs and recovered in a distilled student — are improbable under independent genesis by construction; a 2025 study demonstrates the inheritance itself while arguing that inheritance is fragile and defeatable by paraphrase or inference-time attack [Pan et al., "Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?," ACL 2025, arXiv:2502.11598 — cited here for the premise it assumes and measures, not the robustness conclusion it argues against]; that fragility is a caveat on durability, not on the improbability argument, which is the clean case for this channel. Naturally-occurring idiosyncrasies — formatting habits, structural tics, the glyph-level usus scribendi -- the standard term for a scribe's characteristic writing habits, used for attribution and dating [Reynolds & Wilson, Scribes and Scholars: A Guide to the Transmission of Greek and Latin Literature, 4th ed., Oxford UP, 2013, orig. 1968 -- a canonical general account of scribal transmission, cited for that; whether it specifically treats usus scribendi as a technical concept is not confirmed here, and the more precise candidate anchors for the term itself, per a philological review of this draft, are Pasquali, Storia della tradizione e critica del testo, Le Monnier, 1934, and Bischoff, Latin Palaeography: Antiquity and the Middle Ages, trans. Ó Cróinín & Ganz, Cambridge UP, 1990 -- neither yet confirmed against the primary text [verify]; a 2019 digital-humanities project page using the term in the same sense is not cited as the primary source] — require the same per-item scrutiny errors do, and the attribution literature so far measures separability between models, not the rate at which a given habit would arise under independent training. Distillation detection's reference-based methods — comparing a student's outputs against a named candidate teacher, rather than classifying among a closed set of known systems [Rawat, Chen, Anand, Duan, Rotsted & Min, "Reference-Based Distillation Detection in LLMs," arXiv:2607.09692] — are the closer instrument for the improbability question than attribution accuracy alone, and are what §7.1's validation experiment should lean on.

The corrected inference rule, which this paper offers as its central theoretical claim:

In manuscripts, ancestry is carried by significant shared errors. In models, ancestry is carried by significant shared idiosyncrasies — "significant" bearing the same weight it always did: an idiosyncrasy counts only where independent genesis is improbable, exactly as an error does, and not every stable, discriminating habit clears that bar. Shared errors are predominantly convergent attractors; on at least one measured case, so are shared idiosyncrasies. The stemmatic engine transfers if and only if the improbability-under-independence condition — which philology always carried and the previous draft dropped — is restored and checked per character, not granted to a channel wholesale.

Two corollaries. First, the objection that felled the previous draft is itself the homoplasy/polygenesis problem — the central, named, partially-solved problem of both source disciplines — a fact consistent with the mapping tracking the right structure — though we note, against our own tendency, that reading a refutation's arrival as support has an every-outcome-confirms shape (§11.13), so we state it as consistency, not evidence. Second, the attractor findings carry an operational warning for multi-model review: if capability increases error correlation, then vendor diversity buys less decorrelation every year — the moat erodes on the error channel, and independence must be measured rather than assumed (§7.2).

5. Aggregation: a conditional result, stated symmetrically

The previous draft asserted that majority-vote aggregation of model reviewers is "the wrong aggregation function." Reviewers correctly objected that the supporting mechanisms establish failure conditions, not universal inferiority, and that a defensible case for voting — precision, false-positive suppression, triage economics under low prevalence, Condorcet-style gains under genuine independence — went unaddressed. We restate the claim in its defensible form.

The conditional claim. Majority-vote and consensus aggregation are unsafe where reviewer independence is unmeasured or failing, for two reasons: correlated agreement is duplicated evidence that vote-counting misreads as confirmation, and thresholding discards minority-true findings precisely in the cases where a correlated majority inherits the same blind spot. Both mechanisms are now externally evidenced rather than internally anecdotal: error correlation across providers is measured and rising with capability. That consensus selection thereby filters out minority-correct solutions is our own inference from that association, not a separately measured finding (§4 states this explicitly and the qualifier is not repeated loosely here).

On the limits of this conditional. The conditional above says when aggregation is unsafe as a truth procedure. It does not tell you, in general, which of a practice's own decisions are truth procedures and which are operational choices, and we no longer claim it does. Across three revisions of this paper we attempted a general criterion for that boundary — grounding it first in the formal literature on judgment aggregation (a misuse: that literature is defined by the objects aggregated, not by whether they track truth), then in Condorcet's Jury Theorem with a Popperian test applied recursively to a decision's premises. Each version was repaired by external review; the third still had a hole where it tried to separate observation from aggregation, since almost any pooled judgment can be redescribed as an observation of the pooling. We withdraw the attempt rather than ship a fourth iteration of it, and record the withdrawal here because a paper arguing that fluent overclaiming is the failure mode should not quietly retire its own failed general theory.

What survives is narrower and is all the conditional needs. Where a decision's justification depends on a factual premise that was itself settled by pooling judgments whose independence was unmeasured or failing — the conditional's own trigger, not a narrower one — the independence condition applies to that premise, whatever the decision as a whole is called; a label at the top does not exempt what sits underneath it. Where no such premise is load-bearing — the counterfactual test being whether the decision would change if the premise were false — the condition does not apply. That rule is usable without a general theory of the boundary, and §9 reports a case in this practice's own history where we ignored it and had to withdraw a claim as a result.

Why the condition bites in practice. The condition would be idle if independence were the norm. The measured situation is the opposite: the diversity assumption — that differing data, architecture, and provider produce decorrelated errors — does not hold in practice [Kim, Garg, Peng & Garg, ICML 2025]. Systems we sampled in the contemporary multi-model review space (e.g. a multi-LLM consensus reviewer aggregating across providers; convergence-loop reviewers majority-voting findings and discarding singletons; academic multi-review pipelines with aggregator models) - named concretely, because an unnamed opponent is not criticisable, and named honestly, because naming three systems does not mean all three exemplify the same mechanism. Two cross-vendor systems sit in the configuration the conditional condemns — voting combined with unmeasured independence: Mozilla.ai's Star Chamber queries Claude, GPT and Gemini as separate providers and tiers findings as Consensus (all), Majority (two or more), or Individual observation [Wilson, "The Star Chamber: Multi-LLM Consensus for Code Quality," Mozilla.ai, 5 Mar 2026]; and an academic pipeline aggregates multiple distinct LLMs through an aggregator model, reporting up to +43.67% F1 over single-pass review [arXiv:2509.01494, 2025]. Neither, on inspection, actually discards singleton findings — Star Chamber demotes them to a labelled tier rather than dropping them, which is materially better than discarding, and the aggregator pipeline's mechanism is not described in those terms by its own source. The one system we can cite that does discard singletons via majority vote — Cursor's BugBot, whose own documentation states "majority voting to filter out bugs found during only one pass" [Cursor, "Building a Better Bugbot"] — is not a cross-vendor case: it runs eight parallel passes on one model, which under §3's corrected witness model is legitimate characterisation of a single stochastic witness by repeat sampling, and the conditional does not address it. The discard-singletons mechanism the conditional was written against therefore currently has no named cross-vendor exemplar; we state that gap rather than paper over it with an example that does not quite fit. We do not claim these systems are wrong for their operating points — precision-first triage under alert fatigue is a legitimate loss function — we claim their aggregation is unsafe as a truth procedure under measured correlation, and that the distinction between triage and truth-reconstruction should be explicit in their claims.

The honest cost of the adjudicative alternative. The practice this paper reports uses adjudication: findings are never averaged; a blocking finding blocks; a finding closes only by its raising seat's withdrawal or by demonstration against the artifact. Four separately-commissioned reviewers of the previous draft converged, unprompted, on the same critique of this design, and they are right about its exposure: it grants a de-facto veto to any noisy, stubborn, or adversarial seat; it has no intrinsic false-positive bound, no stalemate rule for oracle-free disputes, and no cost model; it rewards over-production, since findings that cannot be outvoted will be over-raised; it relocates final authority to a human adjudicator who is himself an unanalysed trusted base (§11); and — the consequence §3's corrected witness-unit forces — a blocking finding is itself one draw from a stochastic witness, so the blocking rule inherits sampling noise: a block is treated as a claim to be established against the artifact, never as a verdict, which is exactly the cost the adjudication budget must price. Its viable operating region is low-volume, high-stakes review with oracle access and bounded panels — it does not scale to high-volume triage, and we withdraw any implication that it should. The previous draft's claim that a lone dissenting witness "by construction" preserves truth was circular and is withdrawn; the defensible statement is that under correlated majorities, only a procedure that retains and interrogates dissent can recover minority-true findings — whether a given dissent is one is what adjudication (or an oracle) must establish, at a cost that must be priced, not presumed.

An illustration — explicitly illustration, not evidence, from a single unaudited internal record: an out-of-lineage reviewer once surfaced fifteen issues, one high-severity, that four prior same-lineage rounds had missed. We report it because it motivated the design; the external citations above, not this anecdote, carry the section.

6. The tradition becomes real: corpus recursion

Everything above concerns the synchronic ecology — models as they stand. The paper's strongest claim is diachronic. Until roughly 2022, models were terminal readers of the human archive: leaves on the stemma, copying from everything, contributing nothing back. That has ended. Model output re-enters training corpora through four channels, each a scribal practice under a new name:

  1. Direct distillation — copying from a chosen exemplar; inheritance experimentally demonstrated via watermark persistence into students (§4; Pan et al. 2025).
  2. Ambient corpus contamination — unlabelled machine text scraped into the next generation's corpus: contaminatio as the ecosystem's default condition rather than its accident. Estimates of the synthetic share diverge by method and should not be quoted as settled - a keyword-frequency working paper puts it near 30-40% [Spennemann, arXiv:2504.08755, 2025, not peer-reviewed], while a detector-based sample of 65,000 Common Crawl URLs reports a rise from ~2% in 2020 to roughly half by 2024-25 [Graphite/Originality.ai analysis, reported Oct 2025]; AI-detector reliability is itself disputed.
  3. Deliberate synthetic data — a tradition copying itself. Labs disclose this directly: Meta's Llama 3.1 card records "over 25M synthetically generated examples" in fine-tuning; Microsoft's Phi-3 report describes a corpus of filtered web data plus synthetic LLM-generated data [arXiv:2404.14219]; Anthropic's Claude 3 card lists "data we generate internally." That curation is the only control on this channel is our characterisation, not a claim any of them makes.
  4. Human-mediated diffusion — model phrasings absorbed by human writers and returned to the archive in human text; the documented post-2022 rise of model-characteristic vocabulary in scientific abstracts is this channel measured, at 13.5% of 2024 biomedical abstracts showing LLM-processing signal and higher in some subcorpora [Kobak, González-Márquez, Horvát & Lause, "Delving into LLM-assisted writing in biomedical publications through excess vocabulary," Science Advances, 2025, arXiv:2406.07016].

This channel and §4's attractor argument use the same evidence for different claims, and the distinction is load-bearing. §4 cites lexical overrepresentation ("delve") to show that an idiosyncrasy shared between models need not indicate lineage, because a common training methodology can produce it convergently. This channel cites the same vocabulary's appearance in human text to show a model-to-human transmission path. The two are compatible and separately evidenced: convergence explains why several unrelated models share the habit; diffusion explains how the habit left the models and entered the archive. Neither licenses the other — diffusion into human writing is not evidence that the inter-model sharing was inherited, and the convergent origin does not weaken the diffusion measurement. Conflating them would let a measured transmission channel silently vouch for an unmeasured inheritance claim, which is the substitution §4 exists to block.

Four consequences, each with an exact philological twin:

Stratigraphy. Each model generation deposits dateable idiosyncrasies; models trained on later crawls inherit them. This is dating by shared innovation — and it means lineage inference gets easier as the tradition deepens. The ecology is depositing, unasked, the structure stemmatics needs.

Banalisation. Recursive training on model output erodes distributional tails — rare constructions vanish, mass concentrates on the probable [Shumailov, Shumaylov, Zhao, Papernot, Anderson & Gal, "AI models collapse when trained on recursively generated data," Nature 631:755, 2024; earlier as "The Curse of Recursion," arXiv:2305.17493]. Two qualifications, the second of which is not Shumailov et al.'s own and should not be read as theirs: their strongest results come from pure-replacement recursion, where each generation trains only on its predecessor's output; and where synthetic data accumulates alongside real data rather than replacing it, test error stays bounded and collapse does not occur [Gerstgrasser, Schaeffer, Dey, Rafailov et al., "Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data," arXiv:2404.01413, 2024]. Scribal lectio facilior is the same operation: iterated transmission through a lossy copier that prefers the probable destroys precisely the improbable content. This is where the lectio difficilior row of §3 earns its return, as a law of the tradition rather than a heuristic about samples — and it converts §4's attractors from a mere inference-blocker into the tradition's measured decay mechanism.

The oldest witness. The pre-2022 archive becomes the codex vetustissimus: the one stratum known to predate the contamination event, analogous to pre-nuclear low-background steel — an analogy introduced for this exact purpose days after ChatGPT's release [McDonald, 5 Dec 2022] and since tracked as a project cataloguing pre-contamination corpora [Graham-Cumming, lowbackgroundsteel.ai]. Its value for baselining and dating rises monotonically; data provenance becomes archival science.

The planted mark. This is the channel §4 admits as clean, and the qualification there applies here too: engineered marks are improbable under independent genesis by construction, which is what makes their inheritance evidential, whereas a naturally-occurring habit is not demonstrated to be inherited merely because it is shared. Trap streets and dictionary Mountweazels were deliberately inserted conjunctive errors, planted so a copier could be proven a copier [Agloe, NY — a fictitious hamlet placed by General Drafting Co. in the 1930s as an anagram of its founders' initials; "esquivalience," invented by Christine Lindberg for the New Oxford American Dictionary, 2001, and later found republished elsewhere]. The term of art matters here: a ghost word is an accidental lexicographic error — "dord," a misread annotation that sat in Merriam-Webster's Second New International from 1934 to 1947, caught nobody and proved nothing — whereas a Mountweazel is planted deliberately, which is the only version that functions as a copying-detector. Watermarking and antidistillation fingerprinting are the identical design — synthetic Leitfehler — and their demonstrated inheritance is the cleanest existing proof that the descent channel is real.

The security consequence. Code is part of the corpus. A defect idiom emitted by a model in year N is committed to public repositories, scraped, and trained into year N+2 models. The inheritance channel is demonstrated for planted marks — watermarks, improbable under independent genesis by construction. Whether it extends to naturally-occurring characteristic habits is exactly the question §4 leaves open, not a second demonstrated case; treat it as unresolved on the current evidence, not as a second Leitfehler-grade example. Its further extension to defect idioms is an inference from the same mechanism, stated as such and testable by retrospective tracing — if it holds, vulnerability patterns can be inherited through the archive, not merely regenerated; and the channel is targetable: seeding the corpus with a subtly defective idiom poisons the exemplar a generation of models will copy [inference from the demonstrated channel; no direct study yet — verify against any emerging literature].

Code is also, and this matters for feasibility, the best-conserved textual domain in the tradition: version control gives every serious repository a hash-chained, signed, timestamped manuscript apparatus of its own. The pre-2022 state of a codebase is not a lost archetype but a checkout; dating is a diff against a signed tag. If prospective conservation is winnable anywhere, it is here — and prospective is the operative word: unlike every philologist in history, who reconstructed from the far end of a millennium of loss, this tradition can be instrumented from its origin, with hashes and sealed witnesses from the founding stratum. The defence is conservation — provenance, dating, collation against the oldest witnesses. Any empirical study of this channel must trace, never seed: deliberately planting defects in public corpora is an attack, not an experiment (ethics statement, §13).

Caveats carried with the section: the synthetic fraction of current crawls is uncertain and contested; frontier labs filter aggressively; nothing here is an agency claim about models — transmission structure, not intention.

The time-indexed thesis follows. The reviewers of our previous draft were right about the ecology their objections described — the error channel, where attractors dominate today (§4), and the terminal-reader era whose corpus shaped every present model: against that ecology there is no tradition to reconstruct and the Latin is costume. The recursion is young; the inherited fraction is small now and growing, which is precisely what makes the thesis time-indexed rather than wrong-then-right. The recursion falsifies each objection forward: the archetype (uncontaminated content) becomes practically lost as slop displaces it; the stemma becomes buildable because descent becomes real; the inherited fraction of shared features grows along a proven channel. Stemmatics exists only because traditions corrupt. A tradition has begun to corrupt. That — not any identity claim — is the case for agentic stemmatics, and it is a case that strengthens with time, which is an unusual and falsifiable property for a framework to have (§11 states the falsifier).

7. The program: the operations, specified

The previous draft was justly criticised for performing no distinctively stemmatic operation. This section specifies them as executable work with available ground truth. None is claimed as done.

7.1 The stemma of the models. This experiment performs recensio -- tracing lineage from extant outputs, not conjecturing a reading no witness preserves; what that excludes is stated at §11.7. Reconstruct the descent tree of a model population from output evidence alone — conjunctive idiosyncrasies, per §4's corrected rule — and validate against known lineages: base-model → fine-tune families, teacher → distilled students, checkpoint successions, where ground truth is a matter of record. On present instruments, per §4's split, this experiment has exactly one clean character class today: planted-mark tracing plus reference-based teacher tests against known candidate lineages (distillation detection, teacher identification — Rawat et al., arXiv:2607.09692) — legitimately a fingerprint-phylogeny benchmark on its own terms. The distinctively stemmatic operations this section was criticised for never performing — collation across witnesses, conjunctive grouping of natural characters — become operable only once per-item improbability for natural habits is itself measurable, which it is not yet. The full reconstructed-and-validated stemma appears open [verify: absence claim from a sampled search, scope declared]. Failure modes are informative philology: multi-teacher distillation and synthetic-data admixture are contaminatio formalised, and measuring where reconstruction breaks is part of the result. The experiment is cheap relative to its evidential value, requires no weights access for the subject models' outputs, and its negative result would materially damage this paper's thesis — as a test should.

7.2 The independence metric for review panels. Replace assumed vendor-independence with measured seat-independence: per-question-class error correlation and idiosyncrasy distance between panel seats, tracked over time. This addresses the reviewers' finding that our previous "response-envelope" identity check measures product identity, not error covariance — the envelope tells you who answered; only measured covariance tells you how independent the answers are. The metric is proposed, not built; its own error characteristics would need the same scrutiny as any instrument in §8, and the deeper Bédier limit (§11) is narrowed by measurement, not closed.

7.3 Conservation: the crypt and the degradation clause. Frozen model weights are sealed witnesses: a model trained by year N cannot inherit year-N+1 contamination. An archive of open-weight vintages (closed models cannot be archived by third parties — an asymmetry worth stating: the open-weight ecosystem is the archivable tradition) supports three instruments: temporal decorrelation for panels (a sealed vintage cannot inherit post-seal contamination — though convergent attractors shared before the seal remain shared, so the decorrelation is on the inheritance channel only, per §4); normalisation-drift probes (an idiom that entered the tradition in year N reads as normal to year-N+1 models but remains high-perplexity to earlier, sealed witnesses — a dating instrument); and drift baselines that make corpus-degradation claims measurable rather than atmospheric. For operating panels we adopt a four-point degradation clause: (1) temporal stratification of seats; (2) the independence metric of §7.2, monitored, never assumed; (3) an oracle-share ratchet — the fraction of review weight resting on the mechanical oracles (execution, proofs, tests; the normative-text case of §3 is deliberately excluded here, since a ratchet toward document-reading is not the migration this clause intends) may only increase; (4) out-of-tradition witnesses (human specialists, formal methods) mandatory at the highest tiers. A method that measures its own decay can manage it; only the unmeasured version degrades silently.

8. Field record: instruments that failed, and what the failures evidence

Our operational practice maintains a catalogue of recurring instrument defects — ways a verification tool reports a false result, predominantly false absence: quiet success read as absence; matches discarded by the reader's own filter; summary output read as hits; a control exercising a different command than the claim; control class differing from corpus class (an open family — the skip rules of an arbitrary search cannot be enumerated in advance); patterns encoding a form the corpus does not use; shell/runtime dialect divergence. Each entry carries reproduced instances. Two are worth reporting here for what they evidence.

An absence-warranting gate that failed by its own catalogue. We built a tool to refuse "not found" claims lacking a known-positive control. An independent reviewer constructed six routes to a false "absent," including an inverted warrant (a quiet-success command reported absence precisely when the target was present); the tool's own five-case self-test had passed over all six, one control having been tuned — by its author — to a string that could not trigger the defect. The tool is quarantined, its defects encoded as permanently failing tests. We report this as a documented instance — one instrument, fully autopsied — of the hazard that verification instruments can inherit the disease they check for; the design rules it motivates (controls must be able to fail; graders graded against data their authors did not construct) are engineering responses to a demonstrated possibility, not conclusions from a measured rate.

A live recurrence during this paper's preparation. The operator asserted a document "does not exist" on the strength of a local-directory search while the document stood on the public web — control class ≠ corpus class, the catalogue's own fifth entry, committed by the catalogue's author in the closing hours of drafting. Absence claims in this paper are meant to carry their searched scope explicitly; where one below does not, that is itself an instance of the catalogue's own failure mode, not an exemption from it, and is corrected on discovery rather than defended.

The arithmetic of the one artifact-backed number, corrected. An author-selected mutation set had reported "5/5 killed — 100%"; an AST-enumerated run on the same module yielded 281 mutants of which 143 survived (50.9%) — a kill rate of 49.1% (138/281). The previous draft mislabelled the surviving count with the kill percentage; two reviewers caught it independently. We retain the observation as a single, artifact-backed illustration of one claim only — that generator-selected metrics are the generator's opinion in executable form — and draw no wider generalisation from n=1.

Two dated instances, kept distinct rather than filed as one: the oracle settling what a panel could not, and, separately, calibrated convergence executed rather than prescribed. A critique package arrived against a sibling instrument of this practice (the pre-build protocol's frozen experiment commitments, not this paper), asserting among other things that the instrument's central novelty claim was false. The claim was checkable: primary-source verification (a standards specification — SARIF; two public issue trackers — sarif-spec#120 and in-toto/attestation#77) substantially settled it in one pass, with edge corrections, where the critique's own three commissioned investigations had only asserted it — a correlation we do not measure and therefore do not claim of them either. This is the first instance, dated and filed rather than hypothetical — an application of the discipline §5's conditional recommends (oracle over panel where an oracle exists), not a claim §5 itself states: an oracle question answered by one citation — two different things, properly: the specification is normative-text consultation, an oracle in §3's sense deciding what the text says and nothing about whether it is right; the two issue trackers are filed public records, observed rather than consulted. Neither settles a claim about the world. Together they outperform three agents of one commissioned run, however convergent. That is the first instance. Second, and separately: the critique itself was then put through a three-lineage external seat round — mutually blind, byte-identical brief, this practice's own prior recommendation deliberately withheld so no seat could inherit it — and a majority of the three lineages rejected that withheld recommendation on textual grounds neither had seen the other reach; the third lineage landed partial, permitting the recommendation only as an explicitly declared stretch of the governing clause [internal record, cited not reproduced — the sibling instrument's own filed record]. The verification pass itself carried four minor inaccuracies (an enum undercounted, a field's optionality inverted, an unconfirmed attribution, an uncounted variant) — found by that same verification pass before use, not concealed — so we say substantially settled, with edge corrections, not settled; a paper about unwarranted claims should not overclaim its own verification. We do not offer this as evidence for this paper's thesis; the practice remains one self-observed commercial operation (limit 10) and the critique concerned a different artifact. We offer it as a dated demonstration that the discipline this paper prescribes — oracle over panel where an oracle exists, convergence weighted by whether it answered a withheld or a stated question — was executed, cost real verification work, and reversed a conclusion two prior internal passes had reached. That is what practising the method looks like, separate from whether the method is right.

Internal observations previously offered with load-bearing intent (rerun label instability; same-family blind-spot anecdotes; capability-level claims about self-clearing) are hereby downgraded to what they are: unaudited motivating observations from a single commercial practice, with the selection, incentive, and reproducibility problems that provenance implies (§11). Where external literature now covers the same ground (§4, §5), the citations carry the weight.

9. The experiment: frozen values, quoted — and the gaps, stated

The previous draft claimed "the experiment is specified." That was false when written. The present status is materially different and precisely bounded, and we quote the commitment record rather than paraphrase it [internal record: PBM-VALUES-V8, frozen 16 Aug 2026 after an eight-round adversarial review — 30 findings, 29 closed by the raising seat, one frozen as a declared residual, none by vote; version chain v7→v8 documented with named record files].

How the freeze itself was governed. The freeze followed a pre-committed stopping rule for the review loop: the commissioner declared, on the record and in advance, that round eight was final and the freeze would occur regardless of that round's outcome, any new findings entering as declared residuals rather than amendments — so the record was not iterated until its reviewer tired. The version chain (v7→v8, one amendment applied under the prior round's "amend-then-freeze" verdict) is documented in named record files.

What is frozen (Values 1–4 of the protocol's commitment annex), enumerated:

Discontinuation is bounded: at most one revision, the failure filed first, re-registration with a new pre-committed threshold; a second failure discontinues; results are filed either way. Amendment of frozen values is barred — an amendment attempt is itself a discontinuation-clause event. Two governance flags recorded against this clause by the protocol's own reviewers remain unadjudicated, and we surface rather than resolve them: whether Values 1–3 are protected at a failure-triggered re-registration is unstated in the texts; and the confirmed residual on the ledger-close event (a late-submission refill surface) has a recorded candidate fix held behind the amendment bar rather than silently applied.

Status update, 17 Aug 2026. The values above are discontinued, not failed — no registered run occurred under them. A critique arrived, was substantially settled with edge corrections (§8's phrase, carried here deliberately rather than upgraded), and was put through a three-lineage external seat round.

A conflation in our own record, corrected rather than defended. An earlier draft of this section used the 2-to-1 majority's agreement to settle what the bounded-revision clause means — whether it keys on a run failure or an instrument failure — and then labelled that use "governance" to exempt it from §5's condition. That was wrong on its own terms: the clause's meaning is a truth-candidate, a descriptive claim about a text that could in principle be checked against the clause's drafting history or contradicted by a fourth reading. A correlated three-lineage majority is not a safe way to settle it, and reading the reasons before counting the votes does not change that — two correlated lineages agreeing after reasoning is still a vote with extra steps, not a different epistemic act. We withdraw the implication that the vote established the clause's meaning; the interpretation stands as an open residual, the third lineage's reading retained rather than outvoted, and nothing here adjudicates between them.

What the decision actually rests on. Something narrower, which does not depend on which reading is right: the instrument's governing rule was contested — an event visible in the split itself, not a finding about the clause — and continuing to operate an instrument whose rule is contested cost more than its remaining value. The three lineages' correlation (shared artifact, shared commissioning frame, plausibly the training-corpus hyparchetype §12 names for this paper's own convergence) is a reason their agreement carries little weight as evidence about the clause; it is not a reason against discontinuing, since that decision never depended on their agreement being correct. The commissioner accepted the discontinuation on those grounds. We no longer offer a general account of when aggregation may be used this way — §5 records why that attempt was withdrawn — and this section is an illustration of the discipline being applied, not a demonstration that it works.

What a reader cannot check here. The disputed clause is described in this section but never quoted verbatim, because the record it belongs to is internal and cited rather than reproduced — the standing limitation §12 and the References both record. The ambiguity is at least visible in the paraphrase given above, which conditions discontinuation on "the failure filed first" and on "a second failure" without anywhere stating whether an accepted critique against the instrument's design counts as a failure at all. But a reader wanting to adjudicate the interpretive dispute independently cannot do so from this text, and should treat our characterisation of it as unaudited. We name this rather than leave it for a reader to discover, in the one place where the non-reproduction gap bears directly on a decision this paper reports as made. The values in this section stand as a closed record of what was frozen and why it was retired, not as an active commitment. A successor instrument, built to a narrower kill-screen standard — externally-validated ground truth, a frozen mechanical matching rule, blind non-author scoring — is a separate commissioning, not a revision of this one. The reviewing lineages converged on discontinuing the original instrument; the successor's specific design is this practice's own response, not a design the seats themselves converged on [record: PBM-DISCONTINUATION-RECORD-2026-08-17.md; the successor's own registration record is still owed — this pointer gains a second target when it exists].

What may not be claimed, per the record's own exclusions: a keep verdict licenses only "this builder-plus-protocol bundle was a better purchase than this baseline under these rules" — not "discipline works," not transfer to other builders, not superiority over unconstrained generation (the baseline is rule-governed). The case study is frozen as non-decisional: it prices, and decides nothing. The registered multi-pair run (N = 4 pairs minimum from a to-be-frozen eight-module pool) does not yet exist: the record states verbatim that the pool "does not exist yet; the skeleton binds its shape and size." The record carries a nine-item open-residuals inventory, including the permanent lineage/discipline confound, the partial blindness of the finder seats, and an expectancy leak from case study to registered run. The four frozen values map onto the protocol annex's required commitments (1)–(4); whether they have been separately lodged with the recorder as a registration act is [verify] — the paper claims the freeze, not the lodging. The transfer assumption this experiment exists to test is itself a standing open finding against the protocol (its D5): our strongest prior evidence is review-brief evidence, and its transfer to generation is unverified.

This is, we submit, what a preregistration-forward posture looks like from inside a practice rather than a laboratory: values frozen and quotable, boundaries of licensed claim frozen with them, and the distance to a real registered run stated in the same breath — because a paper that argued fluent overclaiming is the failure mode, and then overclaimed its own experiment, would be its own counterexample. It was, once. §12 records what that cost.

AI Control. The control agenda develops protocols around untrusted models — trusted monitoring, trusted editing, untrusted monitoring [Greenblatt, Shlegeris, Sachan & Roger, "AI Control: Improving Safety Despite Intentional Subversion," ICML 2024, arXiv:2312.06942] — reaching "assurance structure around an untrusted generator" from an alignment starting point. We reach the same architecture from software assurance and philology. We previously called this "decorrelated corroboration" and withdraw the inference: both lines descend from the shared post-2022 untrusted-LLM discourse, and convergence among correlated intellectual lineages is weak evidence — by our own §4. The convergence is noted; nothing is claimed from it.

Formal-verification-gated generation (proof-checked synthesis; state-funded mathematical assurance programmes - the UK's ARIA "Safeguarded AI," a GBP 59M state programme pairing world models with machine-checked proofs to yield quantitative guarantees) is the oracle-space limit of this program, applicable where machine-checkable specifications exist. Our degradation clause's oracle-ratchet is a deliberate migration toward that limit.

The scanner-era instrument stratum. A family of pre-LLM attestation formats already solves adjacent problems for deterministic reviewers: SARIF (OASIS Committee Specification 2019, OASIS Standard 2020) defines a six-value per-result kind enum — fail ("matched"), pass ("clean"), and notApplicable among them — and an artifact-role vocabulary marking what was scanned, with content hashes attached; Microsoft's BinSkim and the OSS Review Toolkit ship comparable per-target disposition and exclusion-reason machinery in production [BinSkim's --kind filter takes Fail/Pass/Review/Open/NotApplicable/Informational per scanned binary; ORT's PathExcludeReason enum records why each excluded path was excluded - nine reason codes including BUILD_TOOL_OF, TEST_OF, DOCUMENTATION_OF. They classify different things, a result disposition and an exclusion rationale, and are comparable in form rather than identical]; FOSSology has carried per-file human clearing decisions since its open-sourcing in December 2007. These are not this paper's competitors; they are its correct ancestors, and the honest lineage runs through them rather than through any bespoke schema. Their shared assumption is the one this paper's subject breaks: a scanner's intended coverage is deducible from its configuration, so a separate totality proof looked unnecessary. That inference is weaker than it sounds, and this paper's own §8 catalogue supplies the counterexamples — runtime failure, unsupported formats, silent exclusions, early termination, and caching all break the step from configured to examined. Configuration establishes intent, not completion. A reviewer whose coverage is not deducible from anything — a human skimming, or a model lineage whose attention over a tree is not observable to the consumer of its output, and whether it is observable to the model itself is the open interpretability question §1 leaves open, not one foreclosed here — has no analogous instrument. The in-toto attestation framework's own maintainers have asked for one publicly, in the framework's attestation repository: issue #77, requesting a generalized human-review predicate, has been open since 2021-12-05, with review-coverage chaining across a diff history named explicitly as the unresolved design question. This paper's collation apparatus — closure as a verification property rather than a progress bar, a property no section of this paper yet specifies as an artifact; §7.2 proposes an independence metric and §9 records a ledger-close rule belonging to a discontinued instrument, neither of which is the verifier meant here — is positioned into that stratum, not against it: what this paper contributes at present is the requirement the gap names, not an instantiated verifier; the intended packaging, if built, would emit SARIF results and an in-toto predicate rather than inventing new interchange, and would claim only the closure verifier and the judgment predicate as new. No such proposal has been submitted; disclosure and timing are separate decisions this paper does not make [verify: pointer to the submission record, if and when it exists].

Multi-model review systems — consensus aggregators, convergence-loop reviewers, panel pipelines, named with their documented aggregation behaviour in §5 — are the nearest mechanism-level prior art and the direct addressee of §5's conditional. Mechanistic interpretability is the complementary science, and the relationship is best stated as a division of labour: interpretability is the physiology of the organism — weights-access required, jurisdiction ending at each lab's own API, coverage presently, by its own practitioners' account, a small fraction of the computation. We found no published figure quantifying that share — an absence claim from a sampled search of the interpretability literature and lab writeups, not an exhaustive one, and scoped here per §8's own requirement rather than asserted flat; the further claim sometimes made, that exhaustive feature extraction would cost more than training compute, we have seen stated by practitioners but not seen quantified, and do not rely on it; agentic stemmatics is the epidemiology of the textual population — outputs only, cross-vendor by construction, its raw material compounding as the tradition deepens (its instruments, per §7, proposed rather than proven). The asymmetries favouring the epidemiological program for the output- warrant question are: jurisdiction (no lab can open a competitor's vat; the tradition is only studiable as a tradition); scaling sign (the interpretability gap plausibly widens with model scale, while collation's raw material compounds); institutional admissibility (a regulator can re-run a collation; no outsider can re-run a feature analysis on closed weights); and historical sufficiency (philology reconstructed the classical canon with a purely behavioural error-model of the scribe — mechanism-knowledge was never a prerequisite for warrant-knowledge). The precedent is epidemiological: the Broad Street pump handle came off decades before the organism was identified [Snow removed the Broad Street pump handle in 1854 without knowing the organism; Koch isolated Vibrio cholerae in 1883 - twenty-nine years later]. Where interpretability is strictly stronger, stated fully because the comparison is propaganda without it: behind unanimous inherited corruption (the perfect hyparchetype leaves no dissenting witness — only mechanism or an oracle sees behind unanimity); on agent-intent questions (deception, sandbagging) that may leave no cross-witness trace; in causal intervention and training-time repair; and prospectively, in mechanism-grade absence claims, which sampling witnesses can never deliver. The unfaithful chain-of-thought result — the scribe's self-commentary diverging from his computation — is an interpretability finding this practice depends on (Turpin et al. 2023; Lanham et al. 2023); the flow runs both ways, and the division of labour has a direction: collation is designed to be the cheap, continuous, ecology-wide triage that tells the expensive microscope where to look — a division this program proposes and has not yet operated at scale. One further relationship is worth stating plainly: the collation records this program would accumulate — dated strata, provenance manifests, sealed vintages, per-question verdict archives — are candidate validation ground truth for future mechanism-level work, the role industrial practice has historically played for the sciences that later explained it - Pasteur's fermentation studies were commissioned out of the French wine and brewing industries' spoilage losses, and the microbiology came out of the practice rather than the reverse [Pasteur, Etudes sur le Vin, 1866; Etudes sur la Biere, 1876].

11. Limits

Stated as the reviewers of the previous draft forced them to be stated, plus those the corrections introduced.

  1. No archetype for judgment questions. Code truth is often plural and specification-relative; reviewer disagreement may reflect different legitimate questions, not corruption of one truth. The framework's reconstruction language applies to decidable questions; elsewhere the panel samples judgment, and calling that reconstruction would be the overreach this draft removed.
  2. Oracle availability bounds the domain. Where tests, types, proofs, execution, or (per §3) the content of a normative text decide, the witness apparatus is subordinate. The framework governs the (large) residue.
  3. Witnesses are stochastic. A lineage must be sampled to be characterised; single-draw verdicts are samples. This burden has no philological precedent and makes several classical operations (e.g. clean eliminatio) unresolvable at the sample level.
  4. The trusted base, located. Verification relocates trust; here it lands on: the human adjudicator (error rate unmeasured, capture and fatigue unmodelled, commercially conflicted in our own practice and declared as such); the independence metric of §7.2 (proposed, unbuilt; its two predecessors failed differently — an early statistical test was formulated backwards, and the later envelope check measured product identity rather than error covariance); the known-positive controls of §8 (author-built, with a documented instance of author-tuning); and the oracles. Who grades the graders is answered only one level up, never finally.
  5. Independence cannot be fully verified from inside — Bédier's critique, extended here by structural analogy rather than claimed as his own finding: his statistical objection concerned stemma-counting specifically, and the transposition to panel independence is ours, not his. Measurement (§7.2) narrows it; nothing closes it.
  6. The correlation unit may be finer than vendor lineage. Prompts, tools, briefs, evidence bundles, and role framing correlate reviewers across vendors; "shape diversity" may dominate brand diversity. Our own review process is a qualitative instance of it, not a measurement: findings that answered our brief's questions tracked commissioning rather than vendor, and we weight convergence accordingly, without claiming a number for an effect we only observed (§12).
  7. Unanimous silence. The method selects among emitted findings; it cannot emend — it will never surface a defect no witness raises. Stemmatic emendatio has no analogue here; the perfect hyparchetype is invisible (§10 assigns this case to interpretability and oracles).
  8. The adversarial hyparchetype. The shared corpus is not only an epistemic limit but an attack surface: corpus poisoning makes correlated failure attacker-controllable (§6). A security-assurance framework must say this of itself.
  9. Prevalence-regime dependence. Dissent-preserving adjudication dominates only where true minority findings exist at meaningful rates; in low-prevalence high-noise regimes it is a false-positive engine. The operating region is stated in §5 and is not universal.
  10. Selection and provenance of the field record. One commercial practice, self-observed, confidential where clients are involved: publication bias, incentive alignment, and non-reproducibility are structural, which is why §8 downgrades internal observations and §9 leans on frozen external-checkable commitments.
  11. Publication Goodhart. If this program succeeds publicly, future models train on its methods and artifacts; detectors become optimisation targets; the trap streets must be renewed. The method's own future validity is time-indexed too.
  12. The catalogue's denominator is unknown. §8 lists detected instrument defects; by the framework's own argument, undetected false absences are invisible, so catalogue completeness is unknowable — an inventory, not a bound.
  13. The Popper episode, demoted. An earlier internal epistemology, built on falsificationism, did not survive its own panel review; only structural theses survived. The previous draft called this "the strongest evidence the method functions." That was an every-outcome-confirms inference — panel kills foundation: rigour; panel spares it: rigour — and is withdrawn. We retain the episode as a consistency demonstration only: the practice applies its procedures to its own doctrine and files the results, whatever they are.
  14. The falsifier for the central thesis. The time-indexed claim of §6 fails if, over successive model generations: idiosyncrasy-based lineage reconstruction does not improve against known phylogenies; inherited-mark persistence (watermarks, glyph habits) attenuates to chance under ordinary training practice; or corpus-recursion channels are shown to contribute negligibly to successor-model behaviour. §7.1 is designed to expose the thesis to exactly this failure.

12. This paper's own collation

Version 0.1 was reviewed by four seats of distinct vendor lineages under an identical adversarial brief, with the author's self-assessment appendix stripped from the review copies. All four returned MAJOR-REVISION; none returned SUBMITTABLE or REJECT. Findings were extracted mechanically from the four filed records into a 209-row inventory (findings, must-verify items, and incorporation requirements), archived alongside this draft; this version is the disposition of that inventory, and the raising seats — not the author — hold closure over their rows.

Calibration, applied to our own evidence per limit 6: convergence among the four on questions the brief asked (the identity overclaim, smuggled efficacy, desk-reject risk) is partly correlation-by-commissioning and is weighted as such. The stronger — not unweighted — signal is the brief-independent convergence: two seats independently found that no distinctively stemmatic operation occurred in the draft; two independently caught the same arithmetic mislabel; four independently converged on the adjudication-DoS critique of §5. Even these share the artifact, the genre of reviewer, and plausibly the training-corpus hyparchetype (§11.5–6), so "stronger" is a comparative within one commissioning, not a claim of measured independence. Those findings shaped this version's structure — §4, §7, and the §5 symmetry are their disposition.

We offer this section as evidence for exactly one claim: that the method is applied to its own artifacts, at cost, with records. It does not answer the closed-loop critique — the author commissioned the review, wrote the brief, and integrates the findings; the review reduces self-certification, it does not abolish it (limit 4). A sibling critique, reviewed by three external lineages mutually blind to each other's output, against a different artifact of this practice, named the mechanism precisely: a closure that can be re-prompted, re-contextualized, and re-run until it succeeds measures the author's persistence, not the finding's resolution, unless the record captures every attempt — not merely the one that closed — outside the author's sole control. That finding applies with undiminished force to the four-seat closure this section describes, in a specific and checkable way: the raw seats' final verdicts are filed and byte-identical to what closed the rows — a reader can inspect them — but nothing in the record distinguishes a first-pass MAJOR-REVISION from a fifth attempt at the same prompt, because attempt counts, discarded runs, and prompt deltas were never captured. The gap is in the history behind the artifact, not in the artifact itself. We name it rather than paper over it.

What was promised, what happened, what exists now. The standing-practice fix proposed against the sibling artifact — an append-only log whose entries the author cannot be the sole gatekeeper over — was said to be "adopted going forward for this paper's own revision rounds, starting with whatever round follows this one" at v0.3. Seven revisions followed. None were logged. That is not a delayed rollout; it is the exact persistence-critique this section describes, committed by this section against itself, and it is corrected here by being stated plainly rather than by a further promise. The historical record — every round from v0.1 through v1.1, including the four-seat convergence this section reports above — will never be covered by attempt-level capture. That gap is permanent. A minimal version of the mechanism — an append-only, hash-chained log of seat invocations — was started at v1.1, covering rounds from that point forward only; it does not yet have external custody, meaning no party other than the author holds or can verify a copy, which is the specific condition the original proposal named and which remains unmet. A log the author alone holds is stronger than no log — it is at minimum tamper-evident to a second party who later obtains a copy — but it is not the independent record the proposal described, and calling it that would be the overclaim this whole passage exists to refuse. The difference between a manifesto and a field report is not that a field report keeps every promise; it is that a field report says which ones it broke, for how long, and exactly what replaced them.

A Bitcoin anchor, and precisely what it does and does not establish. On 20 August 2026 the hash of this log's first entry was committed to the Bitcoin blockchain in an OP_RETURN output, transaction 542748138c9889e745c99fbccd268ec16cafc3e53ebb6fe30b1ee44cdd0145e0, confirmed in block 963282 at 09:23:01 UTC. The payload is the ASCII string CPRLOG1: followed by the SHA-256 entry_hash of entry seq 1. Any reader can fetch that transaction, decode the output, and compare it against the log file's own first entry. What this establishes is that this specific hash existed, and was published, no later than that block — so the author cannot retroactively alter entry 1, or the round it records, without the mismatch being detectable by anyone who checks. What it does not establish is external custody, which is the condition the original proposal named and which remains unmet: no party other than the author holds the log's contents, and a hash is not a copy — it proves what a record was, only to someone who already has the record. Nor does it speak to completeness. The anchor covers one entry, and the log holds four: the three added after it were appended on 20 August 2026, reconstructed from filed artifacts rather than captured at the time of the rounds they record. Those entries are marked as retroactive in the log itself, because a line written a day late does not carry the guarantee a contemporaneous one does — the artifact hashes in them are real, the timestamps of their capture are not. They are also unanchored: only entry 1 is on-chain. This is a tamper-evidence improvement over one entry of a four-entry log, not the independent record the commitment described, and the distinction is exactly the kind this paper elsewhere insists on.

Published, and what publication changed. Later the same day the log was made public at github.com/clvstra/agentic-stemmatics-closure-log, together with a standalone chain verifier and the anchor record. The branch forbids force-pushes and deletions, with the rule applied to administrators including the author; that was confirmed by attempting a history rewrite and having the server reject it, not by reading the setting back — a control nobody has tried to defeat is a claim, not a control. Commits are signed. A reader can therefore clone the log, recompute the chain, check entry 1 against the Bitcoin transaction, and read the commit history for evidence of later editing, none of which requires asking us for anything. What it still is not: the repository sits under the author's own account, the second holder is the practice's other principal, and GitHub is itself a trusted party to the arrangement. None of those three is independent of this practice. The gain is narrower than "external custody" and worth stating exactly — silent revision now requires defeating a public append-only history rather than editing a private file, and the difference is checkable by strangers. Whether that suffices is not ours to declare: the finding was raised by an external seat, and under this section's own rule a row closes only by the seat that raised it. It is recorded here as addressed, and left open pending that disposition.

13. Conclusion, claims, and ethics

We claim: a corrected inference rule for lineage in model populations (idiosyncrasies, not errors, under the improbability condition); a conditional aggregation result grounded in external measurement, with our own alternative's costs stated; a diachronic argument that the model ecology is becoming a textual tradition through measured channels, making the stemmatic frame progressively exact rather than presently identical; an executable validation experiment with ground truth (§7.1) and a falsifier (§11.14); and a conservation program whose value is contingent on that thesis.

We do not claim: that the method produces better code (frozen experiment, discontinued before any run — §9); that adjudication should replace voting outside its stated operating region; that any internal anecdote establishes a rate; or that this paper's survival of its own review process evidences its truth. A preproof becomes a proof only by surviving collation it did not commission.

Ethics. The corpus-inheritance channel must be studied by tracing, never seeding: planting defective idioms in public corpora is an attack on the commons this program exists to conserve. Client-derived observations remain confidential and are marked unaudited wherever used. The author's commercial interest in the method's adoption is declared and is a standing reason to distrust the unexternalised parts of this record — which is why the load-bearing sections are built to rest on external literature, frozen quoted commitments, and specified experiments others can run. Whether they fully succeed in doing so is for readers and the raising seats to judge, not for this paragraph to certify. A citation pass has since been run and its results are in the References below. Seven claims carry [verify], named rather than counted: philological term definitions against standard handbooks (§2), scribal-error directionality (§3), the usus scribendi candidate anchors named at v1.8 (§4), the defect-idiom inheritance inference (§6), the absence claim from a sampled search (§7.1), the registration-lodging question (§9), and the submission-record pointer (§10). This count has now been wrong in six consecutive revisions, for four distinct and separately diagnosed reasons — an em-dash-phrased tag not beginning with the literal string "[verify"; a genuine fix that removed one tag while an unrelated miscount added one back by coincidence; a tag whose text wraps across a line break, which every line-based automated search this session ran, including the one that produced the immediately preceding wrong count, could not see; and, this time, a tag a revision genuinely added — its existence stated in that revision's own header — that the named inventory simply never absorbed. Two independent full-text reads caught the third of those; a decorrelated seat, not the author, caught the fourth. The count is not re-asserted as settled; it is stated as the product of the most careful method tried so far, which is itself the paper's own point about what "checked" can honestly mean.


References

Entries carry caps on what may be assumed of them, and no positive verification mark. An earlier draft of this section tagged every entry the author had checked against a primary source. Two external reviewers, asked independently, both recommended removing that tag: a status field only its author can populate is indistinguishable to a reader from a false one, and placing it beside specific figures implied a content check the artifact does not substantiate. The exhibit was inside this very section — a FOSSology date carried the tag and still contradicted the body text by a year. The marks that remain are limits rather than credentials. [SEC] = established from secondary or reference sources only. [PP] = preprint, not peer-reviewed at time of writing. [REPORTED] = rests on press or vendor reporting this paper did not independently audit. An unmarked entry is an ordinary citation, and claims nothing about who checked it.

Correlated error, attribution, and distillation

Kim, E., Garg, A., Peng, K., & Garg, N. (2025). Correlated Errors in Large Language Models. Proceedings of the 42nd ICML, PMLR 267:30038-30066. arXiv:2506.07962.

Kim, D. (2026). Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles. arXiv:2607.20768. [PP]

Sun, M., Yin, Y., Xu, Z., Kolter, J. Z., & Liu, Z. (2025). Idiosyncrasies in Large Language Models. ICML 2025. arXiv:2502.12150. — five-way attribution reported at 97.1%.

Pan, L., Liu, A., Huang, S., Lu, Y., Hu, X., Wen, L., King, I., & Yu, P. S. (2025). Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? ACL 2025 (Main), 2025.acl-long.648. arXiv:2502.11598. — cited for the inheritance premise it measures, not for its robustness conclusion, which argues inheritance is defeatable. (Misattributed to "Zhao, X." at v0.5-v1.0; corrected on independent GLM/direct-fetch verification, both against the ACL Anthology entry.)

Rawat, R., Chen, S., Anand, A., Duan, M., Rotsted, B., & Min, S. Reference-Based Distillation Detection in LLMs. arXiv:2607.09692. [PP] — reference-based membership inference; despite adjacency in earlier drafts of this paper, it concerns no watermarking.

Model collapse and corpus recursion

Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature 631, 755-759. doi:10.1038/s41586-024-07566-y. Earlier: The Curse of Recursion, arXiv:2305.17493 (2023).

Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. [PP] — the accumulation-vs-replacement result; not part of Shumailov et al.'s own claims.

Kobak, D., Gonzalez-Marquez, R., Horvat, E.-A., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances. doi:10.1126/sciadv.adt3813. arXiv:2406.07016.

Juzek, T. S., & Ward, Z. B. (2025). Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models. COLING 2025. arXiv:2412.11385 — tentatively implicates RLHF; leaves the causal mechanism open.

Spennemann, D. H. R. (2025). Delving into: the quantification of AI-generated content on the internet. arXiv:2504.08755. [PP]

Graphite / Originality.ai analysis of 65,000 Common Crawl URLs, 2020-2025, reported Oct 2025. [REPORTED] — detector-based; detector reliability is itself contested.

Meta, Llama 3.1 Model Card; Microsoft, Phi-3 Technical Report, arXiv:2404.14219; Anthropic, Claude 3 Model Card. — cited for disclosed synthetic-data use in training pipelines.

Reasoning faithfulness and interpretability

Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023. arXiv:2305.04388.

Lanham, T., et al. (2023). Measuring Faithfulness in Chain-of-Thought Reasoning. Anthropic.

AI control and formal assurance

Greenblatt, R., Shlegeris, B., Sachan, K., & Roger, F. (2024). AI Control: Improving Safety Despite Intentional Subversion. ICML 2024. arXiv:2312.06942. — source of trusted monitoring, trusted editing, and untrusted monitoring. The "untrusted advice" protocol is a later, separate fellowship writeup and is not part of this paper.

ARIA (UK Advanced Research and Invention Agency). Safeguarded AI programme.

Attestation formats and the scanner stratum

OASIS. Static Analysis Results Interchange Format (SARIF) v2.1.0. Committee Specification 01, 23 July 2019; OASIS Standard, 27 March 2020. — result.kind is a six-value enum: notApplicable, pass, fail, review, open, informational.

sarif-spec issue #120, "Identify files that were scanned" (opened 2018-03-10, resolved into the artifact-role design).

in-toto/attestation issue #77, "Defining a generalized predicate format for 'human reviews' of artifacts" (opened 2021-12-05 by adityasaky; open at time of writing).

Microsoft. BinSkim, docs/UserGuide.md.

OSS Review Toolkit, PathExcludeReason (model/src/main/kotlin/config).

FOSSology, per-file clearing decisions since its open-sourcing, December 2007.

Multi-model review systems

Wilson, P. (2026). The Star Chamber: Multi-LLM Consensus for Code Quality. Mozilla.ai, 5 March 2026.

Cursor. "Building a Better Bugbot." Cursor Blog. Cited for the aggregation mechanism as documented in the source (eight-pass majority voting); a self-published resolution-rate figure elsewhere in that source is not cited here, was not audited, and does no work in this paper's argument.

Benchmarking and Studying the LLM-based Code Review. arXiv:2509.01494 (2025). [PP] — SWR-Bench; Multi-Agg and Self-Agg aggregation variants.

Philology and textual criticism

Maas, P. (1927). Textkritik. English: Textual Criticism, trans. B. Flower, Clarendon Press, 1958. [SEC] — source of Leitfehler.

Timpanaro, S. (1963). La genesi del metodo del Lachmann. English: The Genesis of Lachmann's Method, ed./trans. G. W. Most, University of Chicago Press, 2005. [SEC]

Bedier, J. (1928). On the manuscript tradition of the Lai de l'Ombre. [SEC] — 105 of 110 surveyed stemmata were two-branched.

Cerquiglini, B. (1989). Eloge de la variante. English: In Praise of the Variant, trans. B. Wing, Johns Hopkins University Press, 1999. [SEC]

Roelli, P. (ed.) (2020). Handbook of Stemmatology: History, Methodology, Digital Approaches. De Gruyter. [SEC]

Roos, T., & Heikkila, T. (2009). Evaluating methods for computer-assisted stemmatology using artificial benchmark data sets. [SEC]

Reynolds, L. D., & Wilson, N. G. (2013). Scribes and Scholars: A Guide to the Transmission of Greek and Latin Literature (4th ed.). Oxford University Press. Orig. 1968. [SEC] — canonical general account of scribal transmission; cited for that, not as a confirmed primary source for usus scribendi specifically. (Duxfield's 2019 digital-humanities project page, cited in earlier drafts, uses the term correctly but is not the primary source; replaced on review.)

Pasquali, G. (1934). Storia della tradizione e critica del testo. Le Monnier. [SEC] — named by a philological review of this draft as the more precise candidate anchor for usus scribendi as a technical concept; not yet confirmed against the primary text.

Bischoff, B. (1990). Latin Palaeography: Antiquity and the Middle Ages, trans. D. Ó Cróinín & D. Ganz. Cambridge University Press. Orig. German 1979, 2nd ed. 1986. [SEC] — named by the same review as an alternative candidate anchor for scribal-habit terminology; not yet confirmed.

Historical analogies

Snow, J. (1854), Broad Street cholera investigation; Koch, R. (1883), isolation of Vibrio cholerae. [SEC]

Pasteur, L. (1866). Etudes sur le Vin; (1876) Etudes sur la Biere. [SEC]

McDonald, K. (5 December 2022), first public use of the low-background-steel analogy for pre-AI-contamination data; Graham-Cumming, J., lowbackgroundsteel.ai. [SEC]

Agloe, New York — trap street placed by General Drafting Co. [SEC] — "esquivalience," coined by C. Lindberg for the New Oxford American Dictionary (2001) as a copyright trap. [SEC] — contrast "dord," an accidental ghost word in Merriam-Webster's Second New International (1934-1947). [SEC]

Still unresolved

The following claims in this draft remain uncited and are marked in text: the paper's own internal records (§8, §9, §12), which are cited but not reproduced; the extension of watermark inheritance to defect idioms (§6), explicitly an inference with no direct study; the §7.1 absence claim about a full reconstructed-and-validated model stemma, which rests on a sampled search with declared scope; and the registration-lodging question in §9. These are named here rather than left for a reader to discover.


Appendix A: Revision history

A note on this build. The drafting copy of this paper carries a working-notes section at the end — an internal maintenance list of open review rows, owed work, and venue thinking, marked in the source as delete-before-submission. It is deliberately not published here. Several entries below refer to it as "the Notes"; those references point at that omitted material. Nothing in the argument, the References, or this history depends on it, and the open findings it tracked are also stated in the body where they bear on a claim. The entries below are reproduced unedited, which is why each carries the "NOT SUBMITTED" stamp it had when it was written: they are dated records of what was true at the time, not statements about this preprint.

Draft v1.16 — 20 Aug 2026. Records the Zenodo deposit: concept DOI 10.5281/zenodo.22030516, version DOI 10.5281/zenodo.22030517 for the v1.15 file, CC BY 4.0, published 20 Aug 2026. The deposited PDF is Agentic-Stemmatics-Soons-2026-v1.15.pdf, 678,933 bytes, SHA-256 d70778bad4a6a6ed873f4db148f65bd053fa65b976e17217685a4f262dafcbd8. Zenodo freezes files on publication, so that byte sequence is now fixed and independently checkable against the record. The paper was not posted to arXiv: arXiv requires an endorsement from an existing author in the subject class for submitters without an institutional affiliation, which this practice does not have. That is a gate on visibility, not on priority, and it is recorded here rather than left as an unexplained absence.

Draft v1.15 — 20 Aug 2026 — NOT SUBMITTED. Adds the author's ORCID (0009-0000-3088-6373) to the byline, so the identifier travels with the PDF rather than living only in a hosting platform's metadata. This revision's note was written directly into this appendix rather than at the top of the paper — the practice v1.14 adopted after the changelog had twice re-accumulated in front of the Abstract.

Draft v1.14 — 20 Aug 2026 — NOT SUBMITTED. This revision moves the notes for v1.11, v1.12 and v1.13 into Appendix A, where the rest of the revision history already lives. They had re-accumulated at the top over three revisions, putting a reader's first page of this paper on its changelog rather than its argument — the same drift v1.11 corrected once already. Nothing in those notes was reworded or cut; only their position changed. Appendix A now carries the complete provenance record from v0.1 through v1.13.

Draft v1.13 — 20 Aug 2026 — NOT SUBMITTED. The closure log is now public (github.com/clvstra/agentic-stemmatics-closure-log), with force-pushes and deletions barred for administrators including the author — verified by attempting a rewrite and having it rejected, rather than by trusting the configuration — and with a standalone verifier anyone can run. This addresses the external-custody finding open since v0.4 without closing it: the repository is the author's own account, the second holder is the practice's other principal, and GitHub is a trusted party, so none of the three is independent. Per this paper's own rule the row closes only by the seat that raised it, so it is recorded as addressed and left open. §12 states the gain precisely and refuses the larger claim.

Draft v1.12 — 20 Aug 2026 — NOT SUBMITTED. The closure log's first entry was anchored to the Bitcoin blockchain this revision (txid 54274813...0145e0, block 963282, 20 Aug 2026 09:23:01 UTC), and §12 now records it together with an explicit account of what it does not buy. It does not close the standing external-custody finding, which has been open since v0.4 and is named in §12 and the Notes as still open: a hash proves what a record was, but only to a reader who already holds the record, so no external party has custody of anything. The log had also fallen out of use since v1.1; the two Grok closure rounds reported above were appended this revision, but retroactively — reconstructed from filed artifacts, marked as such in the log, and not anchored. A retroactively written entry is weaker evidence than a contemporaneous one, and the log now says so on its own face rather than presenting four entries as though they were all captured the same way.

Draft v1.11 — 19 Aug 2026 — NOT SUBMITTED. This revision moves the full revision history — previously the first thing a reader hit, before the Abstract — to Appendix A, after the References. Nothing in the history was reworded or cut in the move; only its position changed. See Appendix A for the complete provenance record from v0.1 through v1.10.

Draft v1.10 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. Grok's closure round on v1.9 returned MINOR-REVISION (GROK-PAPER-V19-CLOSURE-2026-08-19.txt, filed by the operator from Grok's pasted return). All seven v1.7 leftovers closed except F-v03-6 (the closure-log external-custody anchor), which the round explicitly confirmed cannot be paid by wording and should stay open — no change made to it. Two new findings on material no prior round had reviewed as live text, both fixed here: (NF-v19-1) the v1.8 usus scribendi* [verify] tag was never folded into §13's named inventory, which kept saying six when the live body carried seven — a sixth consecutive wrong count, caught this time by a decorrelated seat rather than by either of this session's own full-text reads, now corrected to seven with the omission's cause stated; (NF-v19-2) the v1.8 header claimed to apply "the philology-literate human read... owed since v0.4," but the questions document that read was written for states plainly it needs a classicist and that no automated seat can supply it — and the source of the answers this session worked from was never confirmed before that claim was written. Asked directly this session; not yet answered. The v1.8 header is corrected in place (original text kept, correction stated above it) rather than silently rewritten, and the Notes item stays open until the source is confirmed. Grok also caught a citation hygiene error in the v1.9 header (this note, before this correction, cited a wrong date on its own source file) — noted here as the kind of small self-account slip this paper keeps finding in itself.

Draft v1.9 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. Grok's closure round on v1.7 returned MINOR-REVISION (GROK-PAPER-V17-CLOSURE-2026-08-19.txt), closing two Mediums it had left open across three prior rounds (F-v03-4, F-v03-5) and confirming NF-v06-1/2/5 fixed, but finding the v1.7 subtractive cut honestly stated and incompletely executed: three new Lows (NF-v17-1/2/3) where the header still narrated the withdrawn theory in the present tense, §8 pointed at a "§5 stopping line" the cut had removed, and the surviving §5 rule's trigger ("pooling correlated judgments") was narrower than the conditional it was meant to serve ("unmeasured or failing"), which would have let an unmeasured-but-uncorrelated pool through unchecked. All three are fixed here, along with three Lows left open across multiple prior rounds: F-v03-2 (the "§5 asymmetry" label mis-cited §5 and glued two separate dated instances into one — now separated, with the mis-cite removed), F-v03-7 (§9 overstated the PBM reviewing seats as having converged on a successor's design, when they converged only on discontinuing the original instrument — the successor's design is this practice's own), and NF-v06-4 (§10's "unobservable even to itself" closed a question §1 deliberately left open about the model's own internal state, not just the consumer's — a phrase that survived one prior round undetected because it wraps across a line break, invisible to a naive search, which Grok independently confirmed before fixing it). Not fixed, and not fixable by wording:* F-v03-6, the closure log's external-custody anchor — naming a party other than the author to hold or verify a copy is an operational decision this revision cannot make for the practice; it remains open in the Notes for the author, honestly, rather than closed by asserting an anchor that does not exist.

Draft v1.8 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. Correction, added at v1.10: the paragraph below originally said this revision "applies the philology-literate human read... owed since v0.4." That overclaimed. PHILOLOGY-REVIEW-QUESTIONS-2026-08-18.md states plainly it is written for a classicist, textual critic, or codicologist, and that no automated seat can supply the read it asks for; the source of the answers this revision worked from was never confirmed before that sentence was written, and remains unconfirmed as of v1.10. The correct description is: this revision used the questions document as a checklist and applied the ten answers it was given against the live text, verifying what could be independently checked (bibliographic facts for two candidate citations) and leaving what could not be (whether those sources actually treat the term in question) as open [verify] tags. The owed classicist read itself — sourced, credentialed, attributable — remains open. What follows is the original v1.8 note, otherwise unchanged.

Draft v1.8 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. This revision applies the philology-literate human read of §2–§4 that was owed since v0.4 (PHILOLOGY-REVIEW-QUESTIONS-2026-08-18.md), answered against the questions that document put to the reader. Of ten items plus a central question, six were already correctly implemented by v1.7 and needed no change (the central question; items 1, 5, 6, 8, 10) — confirmed against the actual text, not assumed from the reviewer's own "Implemented" label. Four required real edits, applied here: (2) a disclaimer that Maas's own Leitfehler criterion was qualitative (unwahrscheinlich), not the P-notation this paper formalises it with; (3) a sentence locating the homoplasy/synapomorphy apparatus correctly — the logical distinction at the level classical stemmatics already required, the character-state machinery belonging only to the discipline's modern computational wing; (4) an explicit "structural analogy, not direct transposition" flag on the Bédier comparison, in both §2 and limit 5, since Bédier's finding was about stemma-counting specifically and the extension to panel independence is this paper's own; (9) the usus scribendi citation, which asserted Reynolds & Wilson as "the canonical account of the concept" without that specific claim ever having been confirmed — softened to cite Reynolds & Wilson for what is confirmed (a canonical general account of scribal transmission), with Pasquali's Storia della tradizione e critica del testo (Le Monnier, 1934) and Bischoff's Latin Palaeography (Cambridge UP, 1990) named as the reviewer's more precise candidate anchors for the term itself — both bibliographically verified this revision (author, title, publisher, date), neither yet confirmed to actually treat usus scribendi by name, so both carry [verify] rather than being cited as settled. Item 7 (a mechanical sweep of §§2–4 and §6 for "reconstruction" language doing work only emendatio* could license) was performed this revision: every instance was checked against §7.1's and §11.7's existing recensio/emendatio distinction, and none overreaches it — the sweep is closed clean, with no textual change required.

Draft v1.7 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. This revision subtracts rather than adds. Across v1.3, v1.4 and v1.6 this paper attempted a general criterion separating decisions bound by §5's independence condition from operational choices exempt from it; each attempt was repaired by external review, and the third still had a hole. One reviewing seat noted that if the boundary needed a fourth structural iteration, the need itself was the finding. It did. §5's general theory and §9's application of it are cut — roughly eight thousand characters removed, the first net reduction in this document's history — and replaced by a stated withdrawal plus the one operational rule that survived: where a decision's justification depends on a factual premise settled by pooling correlated judgments, the condition binds that premise whatever the decision is called; where no such premise is load-bearing, it does not. §9 keeps the admission (the vote settled nothing about the clause; the discontinuation rests on the rule having been contested, an event rather than a finding) and drops the theory that was defending it. Nothing in §4, §6, or §7 — the paper's actual contributions — is touched. The prompt for this was a question about whether the review rounds had begun circling: they had, in one section, and the honest answer was to stop building there rather than to iterate a fourth time. v1.4 went to two seats independently. DeepSeek closed row I — open since v0.7 and the one this paper repeatedly called the one that mattered — judging the recursive decomposition to have fixed the displacement defect, and closed rows E and F as well. GLM verified the §5 literature live, the one check no text-only seat could run, and returned the sharpest finding of the round: the boundary was resting on the wrong citation. Judgment aggregation, in the formal literature, is defined by the objects aggregated (sets of yes/no judgments over connected propositions) and not by truth-tracking; its impossibility results concern consistency, not voter independence. Our appeal to it named a different axis than the one we needed. Confirmed independently against the primary source before acting, along with one supporting detail of GLM's own that did not hold — its claim that the entry's examples include normative propositions; they are factual and procedural. §5 is rewritten accordingly: the boundary is the older descriptive/prescriptive one, the independence requirement rides on Condorcet's Jury Theorem and the dependent-vote results (with two precisions GLM's check forced — sufficiently low correlation* rather than strict independence, and binary votes rather than pooled estimates), and the judgment-aggregation appeal is withdrawn as a misuse rather than quietly dropped. §5 also now concedes what it is not inventing: Popper demarcates theories rather than sorting judgments, the descriptive/prescriptive line is Hume's, and the recursion is Quine–Duhem's structure with the stopping line in the classical observation-statement role.

Both seats independently reached the same two fixes, which is convergence worth recording: the stopping line must test a premise's content rather than its mode of access (a record of a pooled judgment makes the pooling observable, never the conclusion observational), and §9's claim that correlation strengthened the non-convergence premise had to go. That claim was ours and we liked it; both seats independently identified it as having the every-outcome-confirms shape §11.13 already catalogues — correlated agreement discounted, correlated disagreement upgraded, no model making either direction wrong — and GLM added the arithmetic, that at three seats and one draw each a 2-to-1 split occurs around 38% of the time even on a strongly-peaked reading. Withdrawn to the narrow claim correlation does not weaken the observation. Row B is fixed by broadening §3's oracle definition to include normative-text consultation (both seats chose this over narrowing §8), with the three carry-throughs GLM identified and we had missed: §7.3's ratchet, §11.2's enumeration, and §9's "grading oracle," which was a third sense of the term hiding one section away. §9's own premises are reclassified in the same pass. Rows B (pending confirmation), D, G, J remain; J is now honestly bounded rather than argued, and §9's self-application is demoted from demonstration to illustration because the disputed clause cannot be reproduced. v1.3 went to DeepSeek, which returned MAJOR CONCERNS (narrower than v1.1) and found a real hole in the boundary v1.3 had just inserted to fix an earlier hole. The §5 test as first written — does the judgment forbid some state of the evidence, or only prescribe an action? — scrutinised the speech act and not the premises underneath it, so a governance decision resting on a pooled, correlated factual premise would pass at the top while laundering a truth-claim one level down. That is the same defect the boundary was written to close, displaced rather than removed, and the criterion did not survive contact. §5 then stated the test recursively, with a counterfactual criterion for which premises are load-bearing and an explicit stopping point so it would not collapse into "everything is truth-reconstruction" — this recursive apparatus is exactly what v1.7 later withdrew (see the note above); it is described here in the past tense because it no longer exists in the live text. §9 then applied the decomposition to its own discontinuation and, at the time, was judged to pass it on the merits — notably because its second premise used the panel's disagreement as directly observable data rather than any seat's conclusion as authority, a "correlation strengthens rather than undermines" framing that DeepSeek and GLM later independently flagged as an every-outcome- confirms claim (see below) and that does not appear in live §9. DeepSeek separately confirmed, by its own counterfactual test, that §9's cost/risk justification is genuinely independent of the withdrawn clause-interpretation claim. Two of its other rows are also addressed: the disputed clause is still not quotable (internal record, cited not reproduced) but §9 now says so plainly and points out that the ambiguity is visible in the paraphrase it does give, and §10's unscoped interpretability absence-claim is scoped per §8's own rule. Rows B, D, F, G remain open and are not text-fixable. A third independent count of the [verify] tags landed on six, matching. A consolidated review report (source unattributed at drafting, later confirmed as the operator's own synthesis of a philology questions brief, a governance briefing, and a Qwen-authored memorandum whose bibliography was found substantially degraded and is not itself cited here) proposed sixteen findings; ten are applied this revision after independent checking, not on the report's word. Verified directly: Timpanaro's actual argument (Bentley, Wolf, pre-Lachmannian origins — confirmed against the book's own chapter contents, not just its existence); the Bédier/Cerquiglini conflation (now separated as opposed rather than continuous reactions); usus scribendi's canonical source (Reynolds & Wilson, replacing a blog citation — chapter locator not claimed, since that detail could not be confirmed). Applied on the strength of their own reasoning: the codex descriptus table row renamed rather than stretched past its logical fit; the hyparchetype disclaimer moved to first use; a closing methodological statement on when the philological frame is dropped, not just worn. The consequential fix at the time was §5/§9: a boundary (judgment aggregation requiring independence versus preference/risk aggregation that does not, then grounded in the formal judgment-aggregation literature) replacing the unfalsifiable "governance decision" label DeepSeek's row I correctly attacked, with §9 no longer defending its 2-to-1 majority by relabelling it. That specific grounding was itself later found to be a misuse of the cited literature (GLM; see the v1.5→v1.6 note below) and the whole boundary apparatus was withdrawn at v1.7 (see the note above). What survives into the live text is narrower: §9 withdraws the claim that the vote settled the clause's meaning, leaves that interpretation an open residual, and rests the discontinuation on independent cost/risk grounds that never depended on the vote being right.

Provenance of v1.1→v1.2 (18 Aug 2026): §12 rewritten that revision to state plainly that its own v0.3 closure-log commitment went unbuilt across seven revisions, that the historical record will never be covered retroactively, and that a minimal version now exists — hash-chained, started at v1.1, no external custody yet — logging the invocation of this very DeepSeek round as its first entry. v1.0 corrected the stale Notes-for-author section that a second unattributed review (no lineage, no envelope, no terminal marker; still unidentified when asked a second time) had read as a live "hard block" — the COI/affiliation metadata itself had been correct since v0.7. That same review returned eleven further findings, ten applied. Two seated reviews of v1.0 then followed — Qwen (formalising its own v0.4 rows and v0.7 assessment; verdict CLOSED on both) and GLM (its first full-text pass after two v0.4 verdicts it had itself disclosed as bundle artifacts, withdrawn this round; verdict MINOR-REVISION). GLM's live primary-source verification caught a real attribution error — the watermark-inheritance citation's first author is Pan, not Zhao, independently confirmed against the ACL Anthology entry before correcting it — and a fifth consecutive wrong [verify]-tag count, this time for a third, distinct reason: the tag lives inside a sentence that wraps across a line break, invisible to every line-based automated search this session ran, including the one that produced the immediately preceding wrong count. Two independent full-text reads, not automated recounts, caught it; both had also independently found the correct number where four prior grep-based passes had not. Also applied: GLM's finding that the restructured §5 exemplar list named three systems but left the discard-singletons mechanism with no cross-vendor exemplar, stated as a gap rather than smoothed over; a softened §8 absence-claim commitment, since four instances did not carry the scope §8 promised of them; and, on GLM's own suggestion, a correlation-weighting sentence ported from §12 into §9, pricing the PBM discontinuation majority at the same epistemic grade §12 already applies to this paper's own four-seat convergence — this does not resolve DeepSeek's row I, it prices it honestly rather than leaving it merely disclosed. Earlier fixes carried forward: the §4/§5 minority-filtering contradiction; the §6 watermark-versus-natural-habit correction; the §7.1 scoping sentence; the §10 artifact-claim correction; undefined DeepSeek row letters and "no response envelope" glossed on use; and a Cursor reference repaired after an earlier cleanup script had broken it, with an unaudited figure dropped rather than left as an unused floating claim. Not applied: the review's own commentary and philology notes, since its lineage is still unconfirmed. DeepSeek's v0.7 findings — labelled B, E, F, G, I, J in that round's own numbering, not otherwise defined in this text: an oracle-citation framing question, an overstrong interpretability claim, the unauthenticated review history, the append-only log's completeness gap, the majority-configuration tension in §9, and the unquoted governing clause — remain untouched; all six want a record or measurement outside the author's sole control that redrafting text cannot supply. Provenance of v0.1: four separately-commissioned model seats of distinct vendor lineages — "distinct lineage" is a commissioning fact, not a measured independence claim (§11.6, §12) — (Grok, Codex — driven; GLM, Qwen — operator-carried, no response envelope — no machine-readable record of which served model, token count, or run identity produced the return, only the operator's own filed header saying so) returned a unanimous MAJOR-REVISION; v0.2 was their disposition, machine-verified by nine checker agents against all 209 extracted rows.

Provenance of v0.2→v0.3 (17 Aug 2026): four bounded additions from the PBM critique arc — §8, §10, §12 (new paragraphs) and §9 (status note) — following verification, a three-lineage external seat round, and an informed audit. Full record: PBM-DISCONTINUATION-RECORD-2026-08-17.md, PAPER-V03-STAGED-EDITS-PBM-ARC-2026-08-17.md.

Provenance of v0.3→v0.4 (same day): v0.3 went through its first closure round against six seat-passes — Grok, Codex, DeepSeek (full/delta access) and GLM twice, once driven on a trimmed bundle and once hand-carried with live web access, plus a full-text hand-carried Qwen pass. Verdicts split cleanly by access: every seat given the full disposition text landed MINOR-REVISION or better on the pre-existing argument (Grok, Qwen; Codex closed 13 of 22 old rows outright); GLM's two MAJOR-REVISION verdicts are substantially an artifact of a bundle design that could not supply disposition text, disclosed as such by GLM itself both times (structural, not content). Six concrete, verified fixes applied this revision: (1) §8's "internally-correlated"/"correlated internal passes" language dropped — asserting correlation without measuring it, even of a critique's own investigations, is the exact discipline this paper argues against; (2) §8's "all three lineages independently rejected" corrected to state the actual 2-firm/1-partial split (Qwen's own verdict on the underlying question was PARTIAL, not a rejection — a Grok fact-check against the filed seat record, not a stylistic softening); (3) §8's citations named directly (SARIF, sarif-spec#120, in-toto/attestation#77) rather than by class, and the four verification inaccuracies' provenance attributed to the verification pass itself; (4) §10's in-toto citation names the repository in reader-facing prose, not only in an internal [verify] tag that would not survive to publication; (5) §10's SARIF timeline precised (2019 Committee Specification, 2020 OASIS Standard) and its kind enum given as six values with three named tokens, resolving both the earlier trichotomy imprecision and a citation-repo ambiguity a hand-carried GLM pass caught by live primary-source verification; (6) §12's closure-log commitment gained a self-aware caveat — the anchor itself is unnamed and a further instance of the same problem if chosen unilaterally by the author. Two items explicitly NOT touched this revision, named rather than silently deferred: formal affiliation/correspondence/COI metadata (Qwen's new Finding 27 — needs real values, not placeholder text) and the remaining [verify] tags (a citation-research task at a different scale, now underway — see the v0.4→v0.5 note immediately below for what has been resolved so far).

Provenance of v0.4→v0.5 (18 Aug 2026): the v0.4 closure round returned six verdicts (Codex MAJOR-REVISION; Grok, DeepSeek, Kimi, GLM, Qwen MINOR-REVISION), converging on a short list of small, confirmed defects — a self-contradictory [verify]-tag count (header said 40, notes said 38; three independent counting methods, Grok/Kimi/GLM/Qwen, converged on 35, not yet reconciled into the text below pending the fuller citation pass this note describes), an Abstract sentence that kept the unmeasured-correlation language §8's fix had dropped elsewhere (fixed below), and a §9 sentence that said "was verified" where §8 says "substantially settled, with edge corrections... not settled" (fixed below, matched to §8's own qualifier). One finding was not mechanical: Kimi, reading the paper cold with no framing inherited from the other five seats, found that §4's corrected inference rule applied its own improbability-under-independent-genesis condition to the error channel but not to the idiosyncrasy channel it replaces errors with — the cited 97% attribution figure measures discriminability (can models be told apart), not improbability (would an independent model develop the same idiosyncrasy by chance), and the paper's own §6 discussion of RLHF-driven vocabulary convergence (e.g. "delve") is a real, citable candidate case of exactly the convergent-attractor problem this channel was supposed to avoid. §4 is revised below to apply the improbability condition per idiosyncrasy rather than to the channel wholesale, splitting engineered markers (watermarks — improbable by construction) from naturally-occurring idiosyncrasies (which need the same per-item scrutiny errors get, not yet uniformly available). A parallel six-thread citation-verification pass checked roughly twenty [verify]-tagged claims against primary sources; most held, several needed real correction (not just sourcing) — including a watermark-inheritance citation that had conflated two unrelated papers, one of which argues the opposite of what it was cited for, and a wrong technical term ("dictionary ghost words," which properly denotes an accidental lexicographic error, not the deliberate copyright trap the paper meant — the correct term is "Mountweazel"). The full bibliography and remaining tag resolution are a separate, still-in-progress pass; this revision applies only the items resolved so far plus the three items above.