A field report and a research program.
L. J. Soons Chief Strategy Officer, CounterProof Research Ltd ORCID: 0009-0000-3088-6373 Correspondence: admin@counterproof.io DOI: 10.5281/zenodo.22030516
Competing interests. The author is an officer of CounterProof Research Ltd, a commercial practice that performs and sells adversarial review services using the method this paper describes. The practice is the source of every internal observation reported here, including the field record in §8, the commitment record in §9, and the review process in §12; it stands to benefit from the method's adoption. No external funding supported this work. The disclosure is placed before the argument rather than after it because the paper's own position is that an interest should be stated before it is argued around, and because several of the limitations in §11 — selection bias, publication bias, the unaudited status of the internal record, and the author's role as adjudicator of his own review panels — follow directly from this relationship rather than being incidental to it.
Draft v1.16 — 20 Aug 2026. Deposited as a preprint on Zenodo, CC BY 4.0. The DOI above is the concept DOI and always resolves to the most recent deposited version; the version deposited on 20 Aug 2026 is v1.15, which is textually identical to this one except that it does not carry the DOI (it could not — the identifier did not exist until the deposit completed). Revision history is in Appendix A.
External claims carry [verify]; internal anecdotes are labelled illustrations, not evidence.
The author operates a commercial practice built on the method described — stated first because
the paper argues that interests must be disclosed before they are argued around.
Large language models emit testimony: fluent claims whose warrant lies outside the act of generation. We propose that the right science for such output is the one philology built for corrupt textual transmission — stemmatics — but not by the identity claim an earlier draft of this paper made and four separately-commissioned reviewers correctly returned for major revision. The defensible thesis is time-indexed and program-shaped: (1) the analogy between model lineages and manuscript witnesses was premature as a description of the 2023–24 model ecology, where shared model errors are — on the measured evidence presently available — dominated by convergent attractors rather than inheritance; (2) it is becoming exact, because model output now re-enters training corpora through measurable transmission channels — direct distillation (with reported watermark-inheritance evidence), ambient corpus contamination, deliberate synthetic data, and human-mediated diffusion — so that the ecology is growing a genuine descent tradition; and (3) the genealogical signal in models is carried not by shared errors (philology's classical conjunctive errors) but by shared idiosyncrasies, for which the attribution and distillation-detection literatures report candidate instruments. From this corrected foundation we derive: a conditional aggregation result — majority voting over model reviewers is unsafe specifically where independence is unmeasured or failing, a condition the correlated-errors literature indicates is the present norm and worsening; an executable experiment — reconstructing the stemma of the model ecosystem from output idiosyncrasies and validating it against known model phylogenies; and a conservation program (archived model vintages as sealed witnesses, provenance strata, and a degradation clause for multi-model review panels). We report the failure record of our own instruments, including the reviewed failure of this paper's previous draft, whose disposition table is part of the method, and a second, dated instance from a sibling instrument of this practice — external verification settling a claim that internal agreement alone could not, and its own governance clause exited visibly rather than laundered. Commitment values for the practice's first controlled comparison were frozen and quoted from their commitment record; that instrument has since been discontinued on critique, before any run — a decision we report rather than paper over. We claim a framing, a corrected inference rule, and a program — not demonstrated efficacy.
A conventional program can, in principle, be read; its behaviour derives from its text. A language model's output cannot be read this way. It is a claim — "this code is correct," "no such vulnerability exists" — and the claim's truth is not reliably recoverable from the fluency or internal consistency of the text asserting it. Structurally, the mechanism that emits a well-grounded claim and the mechanism that emits a baseless one are the same mechanism; whether any internal mark distinguishes them is an open question for interpretability, not a property available to the consumer of the output [Turpin, Michael, Perez & Bowman, "Language Models Don't Always Say What They Think," NeurIPS 2023, arXiv:2305.04388; Lanham et al., "Measuring Faithfulness in Chain-of-Thought Reasoning," Anthropic, 2023].
We therefore treat model outputs as witness testimony — the witness proper being the model lineage, sampled repeatedly, of which any single output is one draw (§3) — to be weighed, collated, and attributed rather than trusted or merely averaged. The discipline that industrialised witness-weighing for text is philology, and its genealogical wing — stemmatics — is the specific craft of recovering truth from multiple, individually corrupt, partially dependent witnesses.
What changed since the previous draft. Version 0.1 of this paper asserted the mapping was "structurally identical" and treated stemmatics' inference engine as directly transferable. Four separately-commissioned reviewers unanimously returned that draft for major revision, and their two deepest findings — that the paper never performed a distinctively stemmatic operation, and that shared model errors indicate convergence rather than descent — were correct. This version is built on those findings rather than around them. The result is a weaker claim about the present, a stronger claim about the trajectory, a corrected inference rule, and an experiment with available ground truth.
Contributions. (i) A conditioned mapping between stemmatics and multi-model verification, with its failure points stated as precisely as its correspondences (§3). (ii) The homoplasy correction: why the classical conjunctive-error inference does not transfer to models, and the corrected rule — model ancestry rides shared idiosyncrasies that are individually improbable under independent genesis, not shared errors — with the measured evidence for the attractor half, and an explicit account of which idiosyncrasies do and do not presently clear that bar (§4). (iii) A conditional aggregation result for multi-model review, grounded in the external correlated-errors literature rather than in our internal anecdotes, with the failure modes of our own adjudicative alternative stated symmetrically (§5). (iv) The corpus-recursion argument: the measurable channels by which the model ecology is becoming a textual tradition, and what follows for dating, conservation, and security (§6). (v) An executable program: the stemma-of-models experiment against known lineages; an empirical independence metric for review panels; and a conservation architecture with a degradation clause (§7). (vi) A field record of instrument failures, including this paper's own reviewed failure and a dated external instance from a sibling artifact, offered as evidence for the framing and explicitly not for efficacy (§8, §12). (vii) The frozen commitment record of the practice's first controlled comparison, quoted with its own must-not-claim boundaries, and its subsequent discontinuation, reported with the same discipline (§9).
We do not claim the method produces better code. That claim is reserved for a registered experiment whose commitment values existed and whose run never occurred (§9).
Classical texts survive as copies of copies; the original — the archetype — is lost. The genealogical method conventionally attributed to Karl Lachmann — an attribution Timpanaro qualifies not by disputing that the method exists under that name, but by demonstrating that Lachmann adapted and formalised existing classical philological practice (the humanist tradition of emendatio ope codicum running through Bentley, and F. A. Wolf's Prolegomena) rather than inventing the method ex nihilo [Timpanaro, The Genesis of Lachmann's Method, trans. Most, Univ. of Chicago Press, 2005; orig. La genesi del metodo del Lachmann, 1963 — the contestation is historiographical, about originality, not ontological, about whether the method is real] — reconstructs archetypal readings from the pattern of agreement and disagreement among witnesses. Its engine is the conjunctive error: a shared error so arbitrary that independent double-genesis is improbable, which therefore evidences common ancestry.
Two conditions, present in the discipline from the start, matter for everything that follows.
First, the improbability condition. Not every shared error is conjunctive. Philology explicitly excludes polygenetic errors — banalisations, easy slips, orthographic modernisation — precisely because independent scribes converge on them. Only significant errors (Maas's Leitfehler, coined by analogy to geology's index fossils and subdivided into separative and conjunctive errors) carry genealogical signal [Maas, Textkritik, 1927; English Textual Criticism, trans. Flower, Clarendon, 1958]. The inference was never "shared error ⇒ descent"; it was "shared feature ⇒ descent only where P(shared | independent genesis) is negligible" — Maas's own criterion was qualitative (unwahrscheinlich, improbable, under independent origin); the P-notation is this paper's own formalisation for computational operationalisation, not language Maas himself used. Biology faces the identical problem as homoplasy (convergence) versus synapomorphy (shared derived character). The distinction itself is imported at the logical level, where classical stemmatics already required it, holistically and editor-dependently, without any formal apparatus; the machinery that makes it computable — character-state matrices, cladistic parsimony — belongs to the modern computational wing of the discipline, and computational stemmatology explicitly shares phylogenetics' machinery for it, including benchmark evaluation against artificial ground-truth traditions [Roelli (ed.), Handbook of Stemmatology, De Gruyter, 2020; Roos & Heikkilä, "Evaluating methods for computer-assisted stemmatology using artificial benchmark data sets," 2009].
Second, the discipline's own deflationary history. Stemmatics is not an unchallenged triumph. Bédier observed that reconstructed trees came out suspiciously two-branched and concluded the method partly manufactures the structure it claims to find [Bédier, 1928 — of 110 published stemmata surveyed, 105 were two-branched]. The field's response split into two directions the paper's earlier drafts ran together as one trajectory, and should not: Bédier himself retreated to best-manuscript editing — he still sought a single recoverable text, abandoning only the genealogical method as the instrument for reaching it — while the later New Philology [Cerquiglini, In Praise of the Variant, trans. Wing, Johns Hopkins, 1999, orig. Eloge de la variante, 1989] took the opposite turn, abandoning the goal of a single archetype entirely in favour of variance as the object of study. We cite both as the discipline's own deflationary history, not as one continuous move. We use stemmatics as an engine with its internal critiques attached. The Bédier problem reappears in our setting by structural analogy, not direct transposition — both cases involve a method manufacturing confidence in the very structure it cannot verify from inside — as the impossibility of fully verifying panel independence from inside the panel (§11).
Four further terms recur [definitions against standard handbooks — verify]: codex descriptus
(a copy of an extant witness, eliminated as informationally redundant); contaminatio (a witness
copied from multiple exemplars — mixed ancestry, the hardest corruption to detect); lectio
difficilior potior (the harder reading is likelier original, because copying drifts toward the
easy); and the hyparchetype (a shared corrupt intermediate ancestor, behind which no collation
can see).
A note on the frame's status. We do not claim a 1:1 ontological identity between manuscripts and LLMs. We adopt the inferential logic of stemmatics -- its handling of transmission, corruption, independence, and the improbability condition -- as a formal epistemological framework for machine testimony. Where the classical vocabulary maps precisely onto the computational reality (contamination, stratigraphy, the improbability condition), we use it. Where it does not (strict eliminatio, emendatio), we discard the Latin and keep only the logical operation -- §3's table is revised in several rows on exactly this basis. The value of the philological frame is not its vocabulary; it is a centuries-tested discipline for managing correlated witnesses, and a paper using that discipline should be willing to drop the costume the moment the costume stops fitting.
The table below replaces the previous draft's. Each row now carries its status; three rows that reviewers showed to be false or forced have been corrected or removed, and the corrections are substantive, not cosmetic.
| Stemmatics | Multi-model verification | Status of the row |
|---|---|---|
| Lost archetype | Ground truth about the artifact | Conditioned. Holds only for decidable questions — throughout this paper that names the operational sense, a question some oracle settles — an execution, a type-check, a test, a proof, or consultation of a normative text where the question is what that text says (a specification's enum, an issue's status), which settles the content of the text and never whether the text is correct or whether some claim about the world holds — and not the computability-theoretic sense of the term; the general property is undecidable there, and no argument here needs it to be otherwise. For judgment questions (severity, design soundness) there may be no single archetype — reviewers may answer different legitimate questions, not corrupt one truth. And unlike the philologist's, our archetype is often extant but expensive: the code can be read, tests can be run. Where an oracle exists, collation is subordinate to it. This bounds the framework's domain (§11). |
| Manuscript witness | A model lineage, sampled repeatedly | Corrected. The previous draft mapped the witness to a single output, which its own limits section contradicted: a model returns different verdicts on identical input. The atomic unit is the lineage-as-distribution; a single output is one draw from a witness, not a witness. Consequences for blocking rules are drawn in §5 (a block is one draw, to be established, not a verdict) and §11.3. |
| Conjunctive error ⇒ shared ancestry | ~~Shared model error ⇒ shared training lineage~~ | Replaced. The classical engine does not transfer on the error channel — §4. The corrected rule: significant shared idiosyncrasies ⇒ lineage. |
| Witness redundancy / sampling variance | Second same-family output: discounted for measured correlation, not eliminated | Renamed. Philology's strict eliminatio codicum descriptorum removes a copy because it adds zero genealogical information. A second LLM sample is not that: it adds distributional information about the lineage's variance, so the operation is discounting for measured correlation, not elimination. The logical character of the operation has changed enough that we drop the Latin term for this row rather than stretch it to cover a different operation. What eliminatio correctly warns against — counting correlated witnesses as independent — is a weighting error, and that warning we keep. |
| Contaminatio | Cross-seat contamination: feeding one reviewer another's findings | Holds, narrowly. Illustrated (not evidenced) in our field record (§8); review survival is not offered as warrant (§13). A shared evidence bundle is a milder relative — common inputs rather than mixed exemplars — and the two should not be lumped. |
| Independent agreement ⇒ likely truth | Decorrelated agreement ⇒ likely truth | Conditioned on measured independence — which cannot be assumed from vendor identity (§4, §5) and cannot be fully verified from inside the panel (§11). |
| Lectio difficilior | ~~Distrust the fluent output~~ | Removed as a synchronic rule; re-derived diachronically in §6. As a rule for choosing between two present outputs it is decorative — "difficulty" is not a defined signal, and scribal error is not uniformly directional [verify]. But as a claim about iterated transmission it becomes exact: recursive training on model output erodes distributional tails (§6; Shumailov et al., Nature 631:755, 2024) as scribal banalisation erodes hard readings. The row earns its life at the level of the tradition, not the sample. |
| Hyparchetype | Shared training corpus | Functional analogue, not a strict mapping. Classically a hyparchetype is a specific, lost, reconstructible intermediate manuscript; we extend the term to a shared training corpus, which is not a single ancestor. The extension is deliberate: we use "hyparchetype" for the risk class — correlated blind spots behind which collation cannot see — because that risk class is what matters here, not the entity-status of the common source. We acknowledge this goes beyond the term's strict codicological meaning, stated here rather than left implicit (§6, §11). |
Two admissions govern the table's use. The generic force of several rows is common-mode-failure analysis, which reliability engineering possesses without Latin; what the philological frame adds beyond vocabulary is (a) the diachronic apparatus — descent, strata, contamination, conservation — which becomes literal in §6, and (b) a worked historical case in which output-level science proved adequate for reconstruction-grade warrant without mechanism-level access — in that domain (§10; whether the adequacy transfers here is part of what §7 must show). And the previous draft's most damning review finding — that no distinctively stemmatic operation was ever performed — is accepted and answered not by rewording but by §7, where the operations are specified as experiments.
The deepest objection raised against the previous draft deserves full statement. Scribal conjunctive errors evidence descent because they are improbable coincidences under a copy process. Model errors are not like this: two models trained on overlapping data toward similar objectives will converge on the same plausible-but-wrong answers as attractors — the mode of a shared objective — with no copying involved. Shared falsehood therefore fails to evidence shared ancestry. As stated, this threatened the framework's engine.
The resolution has two measured halves.
The attractor half is real, and externally measured. The correlated-errors literature reports, across hundreds of models, that when two models both err they agree on the same wrong answer at rates far above chance (~60% in leaderboard settings); that error correlation persists across distinct architectures and providers; and — most consequentially — that more capable models exhibit more correlated errors [Kim, Garg, Peng & Garg, "Correlated Errors in Large Language Models," ICML 2025, PMLR 267:30038-30066, arXiv:2506.07962 - over 350 models across two leaderboards and a resume-screening task]. A capability-controlled audit finds majority vote beats the best single ensemble member in under 10% of canonical three-model subsets, an effect its author characterises as modest and configuration-dependent rather than decisive [Kim, "Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles," arXiv:2607.20768, 2026 - preprint, not peer-reviewed at time of writing; 31,900 subsets over 30 models on MMLU-Pro]. That consensus selection thereby filters out minority-correct answers is our inference from the association between shared error overlap and lost gain, not a finding either paper states. So on the error channel, the objection stands: shared errors are predominantly polygenetic. Philology's own exclusion rule, applied honestly, discards most of them as genealogically void.
But the idiosyncrasy channel carries real, measured signal — establishing which idiosyncrasies clear the same bar errors must clear took a second look, prompted by an external reviewer who found the first version of this section had not done it. Models exhibit stable, model-specific idiosyncrasies — word-level distributions, structural habits, formatting signatures — sufficient for 97.1% five-way attribution between major systems, robust to paraphrase [Sun, Yin, Xu, Kolter & Liu, "Idiosyncrasies in Large Language Models," ICML 2025, arXiv:2502.12150]. That figure is discriminability: each model's idiosyncrasy profile differs enough from its neighbours' to classify unseen text. Discriminability is not the improbability-under-independent-genesis condition §2 requires, and at least one prominent idiosyncrasy fails it on present evidence: lexical overrepresentation of words like "delve" is measured across unrelated post-2022 models and, independently, in human scientific writing, with a candidate common cause in shared RLHF methodology rather than shared lineage [Juzek & Ward, "Why Does ChatGPT 'Delve' So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models," COLING 2025, arXiv:2412.11385 — the paper leaves the causal mechanism open, which is why we say candidate]. That is the same convergent-objective attractor §4's first half already uses to void shared errors as evidence, applied here to a habit instead of a mistake — and it means the idiosyncrasy channel inherits the error channel's own problem unless checked per item.
The corrected rule therefore needs the improbability condition applied idiosyncrasy-by-idiosyncrasy,
not granted to the channel wholesale, and the channel splits into two sub-classes that earn it on
different grounds. Engineered markers — a statistical mark planted in a teacher's outputs and
recovered in a distilled student — are improbable under independent genesis by construction; a
2025 study demonstrates the inheritance itself while arguing that inheritance is fragile and
defeatable by paraphrase or inference-time attack [Pan et al., "Can LLM Watermarks Robustly
Prevent Unauthorized Knowledge Distillation?," ACL 2025, arXiv:2502.11598 — cited here for the
premise it assumes and measures, not the robustness conclusion it argues against]; that fragility
is a caveat on durability, not on the improbability argument, which is the clean case for this
channel. Naturally-occurring idiosyncrasies — formatting habits, structural tics, the
glyph-level usus scribendi -- the standard term for a scribe's characteristic writing habits,
used for attribution and dating [Reynolds & Wilson, Scribes and Scholars: A Guide to the
Transmission of Greek and Latin Literature, 4th ed., Oxford UP, 2013, orig. 1968 -- a canonical
general account of scribal transmission, cited for that; whether it specifically treats usus
scribendi as a technical concept is not confirmed here, and the more precise candidate anchors for
the term itself, per a philological review of this draft, are Pasquali, Storia della tradizione e
critica del testo, Le Monnier, 1934, and Bischoff, Latin Palaeography: Antiquity and the Middle
Ages, trans. Ó Cróinín & Ganz, Cambridge UP, 1990 -- neither yet confirmed against the primary
text [verify]; a 2019 digital-humanities project page using the term in the same sense is not
cited as the primary source] — require the same per-item
scrutiny errors do, and the attribution literature so far measures separability between models,
not the rate at which a given habit would arise under independent training. Distillation
detection's reference-based methods — comparing a student's outputs against a named candidate
teacher, rather than classifying among a closed set of known systems [Rawat, Chen, Anand, Duan,
Rotsted & Min, "Reference-Based Distillation Detection in LLMs," arXiv:2607.09692] — are the
closer instrument for the improbability question than attribution accuracy alone, and are what
§7.1's validation experiment should lean on.
The corrected inference rule, which this paper offers as its central theoretical claim:
In manuscripts, ancestry is carried by significant shared errors. In models, ancestry is carried by significant shared idiosyncrasies — "significant" bearing the same weight it always did: an idiosyncrasy counts only where independent genesis is improbable, exactly as an error does, and not every stable, discriminating habit clears that bar. Shared errors are predominantly convergent attractors; on at least one measured case, so are shared idiosyncrasies. The stemmatic engine transfers if and only if the improbability-under-independence condition — which philology always carried and the previous draft dropped — is restored and checked per character, not granted to a channel wholesale.
Two corollaries. First, the objection that felled the previous draft is itself the homoplasy/polygenesis problem — the central, named, partially-solved problem of both source disciplines — a fact consistent with the mapping tracking the right structure — though we note, against our own tendency, that reading a refutation's arrival as support has an every-outcome-confirms shape (§11.13), so we state it as consistency, not evidence. Second, the attractor findings carry an operational warning for multi-model review: if capability increases error correlation, then vendor diversity buys less decorrelation every year — the moat erodes on the error channel, and independence must be measured rather than assumed (§7.2).
The previous draft asserted that majority-vote aggregation of model reviewers is "the wrong aggregation function." Reviewers correctly objected that the supporting mechanisms establish failure conditions, not universal inferiority, and that a defensible case for voting — precision, false-positive suppression, triage economics under low prevalence, Condorcet-style gains under genuine independence — went unaddressed. We restate the claim in its defensible form.
The conditional claim. Majority-vote and consensus aggregation are unsafe where reviewer independence is unmeasured or failing, for two reasons: correlated agreement is duplicated evidence that vote-counting misreads as confirmation, and thresholding discards minority-true findings precisely in the cases where a correlated majority inherits the same blind spot. Both mechanisms are now externally evidenced rather than internally anecdotal: error correlation across providers is measured and rising with capability. That consensus selection thereby filters out minority-correct solutions is our own inference from that association, not a separately measured finding (§4 states this explicitly and the qualifier is not repeated loosely here).
On the limits of this conditional. The conditional above says when aggregation is unsafe as a truth procedure. It does not tell you, in general, which of a practice's own decisions are truth procedures and which are operational choices, and we no longer claim it does. Across three revisions of this paper we attempted a general criterion for that boundary — grounding it first in the formal literature on judgment aggregation (a misuse: that literature is defined by the objects aggregated, not by whether they track truth), then in Condorcet's Jury Theorem with a Popperian test applied recursively to a decision's premises. Each version was repaired by external review; the third still had a hole where it tried to separate observation from aggregation, since almost any pooled judgment can be redescribed as an observation of the pooling. We withdraw the attempt rather than ship a fourth iteration of it, and record the withdrawal here because a paper arguing that fluent overclaiming is the failure mode should not quietly retire its own failed general theory.
What survives is narrower and is all the conditional needs. Where a decision's justification depends on a factual premise that was itself settled by pooling judgments whose independence was unmeasured or failing — the conditional's own trigger, not a narrower one — the independence condition applies to that premise, whatever the decision as a whole is called; a label at the top does not exempt what sits underneath it. Where no such premise is load-bearing — the counterfactual test being whether the decision would change if the premise were false — the condition does not apply. That rule is usable without a general theory of the boundary, and §9 reports a case in this practice's own history where we ignored it and had to withdraw a claim as a result.
Why the condition bites in practice. The condition would be idle if independence were the norm. The measured situation is the opposite: the diversity assumption — that differing data, architecture, and provider produce decorrelated errors — does not hold in practice [Kim, Garg, Peng & Garg, ICML 2025]. Systems we sampled in the contemporary multi-model review space (e.g. a multi-LLM consensus reviewer aggregating across providers; convergence-loop reviewers majority-voting findings and discarding singletons; academic multi-review pipelines with aggregator models) - named concretely, because an unnamed opponent is not criticisable, and named honestly, because naming three systems does not mean all three exemplify the same mechanism. Two cross-vendor systems sit in the configuration the conditional condemns — voting combined with unmeasured independence: Mozilla.ai's Star Chamber queries Claude, GPT and Gemini as separate providers and tiers findings as Consensus (all), Majority (two or more), or Individual observation [Wilson, "The Star Chamber: Multi-LLM Consensus for Code Quality," Mozilla.ai, 5 Mar 2026]; and an academic pipeline aggregates multiple distinct LLMs through an aggregator model, reporting up to +43.67% F1 over single-pass review [arXiv:2509.01494, 2025]. Neither, on inspection, actually discards singleton findings — Star Chamber demotes them to a labelled tier rather than dropping them, which is materially better than discarding, and the aggregator pipeline's mechanism is not described in those terms by its own source. The one system we can cite that does discard singletons via majority vote — Cursor's BugBot, whose own documentation states "majority voting to filter out bugs found during only one pass" [Cursor, "Building a Better Bugbot"] — is not a cross-vendor case: it runs eight parallel passes on one model, which under §3's corrected witness model is legitimate characterisation of a single stochastic witness by repeat sampling, and the conditional does not address it. The discard-singletons mechanism the conditional was written against therefore currently has no named cross-vendor exemplar; we state that gap rather than paper over it with an example that does not quite fit. We do not claim these systems are wrong for their operating points — precision-first triage under alert fatigue is a legitimate loss function — we claim their aggregation is unsafe as a truth procedure under measured correlation, and that the distinction between triage and truth-reconstruction should be explicit in their claims.
The honest cost of the adjudicative alternative. The practice this paper reports uses adjudication: findings are never averaged; a blocking finding blocks; a finding closes only by its raising seat's withdrawal or by demonstration against the artifact. Four separately-commissioned reviewers of the previous draft converged, unprompted, on the same critique of this design, and they are right about its exposure: it grants a de-facto veto to any noisy, stubborn, or adversarial seat; it has no intrinsic false-positive bound, no stalemate rule for oracle-free disputes, and no cost model; it rewards over-production, since findings that cannot be outvoted will be over-raised; it relocates final authority to a human adjudicator who is himself an unanalysed trusted base (§11); and — the consequence §3's corrected witness-unit forces — a blocking finding is itself one draw from a stochastic witness, so the blocking rule inherits sampling noise: a block is treated as a claim to be established against the artifact, never as a verdict, which is exactly the cost the adjudication budget must price. Its viable operating region is low-volume, high-stakes review with oracle access and bounded panels — it does not scale to high-volume triage, and we withdraw any implication that it should. The previous draft's claim that a lone dissenting witness "by construction" preserves truth was circular and is withdrawn; the defensible statement is that under correlated majorities, only a procedure that retains and interrogates dissent can recover minority-true findings — whether a given dissent is one is what adjudication (or an oracle) must establish, at a cost that must be priced, not presumed.
An illustration — explicitly illustration, not evidence, from a single unaudited internal record: an out-of-lineage reviewer once surfaced fifteen issues, one high-severity, that four prior same-lineage rounds had missed. We report it because it motivated the design; the external citations above, not this anecdote, carry the section.
Everything above concerns the synchronic ecology — models as they stand. The paper's strongest claim is diachronic. Until roughly 2022, models were terminal readers of the human archive: leaves on the stemma, copying from everything, contributing nothing back. That has ended. Model output re-enters training corpora through four channels, each a scribal practice under a new name:
This channel and §4's attractor argument use the same evidence for different claims, and the distinction is load-bearing. §4 cites lexical overrepresentation ("delve") to show that an idiosyncrasy shared between models need not indicate lineage, because a common training methodology can produce it convergently. This channel cites the same vocabulary's appearance in human text to show a model-to-human transmission path. The two are compatible and separately evidenced: convergence explains why several unrelated models share the habit; diffusion explains how the habit left the models and entered the archive. Neither licenses the other — diffusion into human writing is not evidence that the inter-model sharing was inherited, and the convergent origin does not weaken the diffusion measurement. Conflating them would let a measured transmission channel silently vouch for an unmeasured inheritance claim, which is the substitution §4 exists to block.
Four consequences, each with an exact philological twin:
Stratigraphy. Each model generation deposits dateable idiosyncrasies; models trained on later crawls inherit them. This is dating by shared innovation — and it means lineage inference gets easier as the tradition deepens. The ecology is depositing, unasked, the structure stemmatics needs.
Banalisation. Recursive training on model output erodes distributional tails — rare constructions vanish, mass concentrates on the probable [Shumailov, Shumaylov, Zhao, Papernot, Anderson & Gal, "AI models collapse when trained on recursively generated data," Nature 631:755, 2024; earlier as "The Curse of Recursion," arXiv:2305.17493]. Two qualifications, the second of which is not Shumailov et al.'s own and should not be read as theirs: their strongest results come from pure-replacement recursion, where each generation trains only on its predecessor's output; and where synthetic data accumulates alongside real data rather than replacing it, test error stays bounded and collapse does not occur [Gerstgrasser, Schaeffer, Dey, Rafailov et al., "Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data," arXiv:2404.01413, 2024]. Scribal lectio facilior is the same operation: iterated transmission through a lossy copier that prefers the probable destroys precisely the improbable content. This is where the lectio difficilior row of §3 earns its return, as a law of the tradition rather than a heuristic about samples — and it converts §4's attractors from a mere inference-blocker into the tradition's measured decay mechanism.
The oldest witness. The pre-2022 archive becomes the codex vetustissimus: the one stratum known to predate the contamination event, analogous to pre-nuclear low-background steel — an analogy introduced for this exact purpose days after ChatGPT's release [McDonald, 5 Dec 2022] and since tracked as a project cataloguing pre-contamination corpora [Graham-Cumming, lowbackgroundsteel.ai]. Its value for baselining and dating rises monotonically; data provenance becomes archival science.
The planted mark. This is the channel §4 admits as clean, and the qualification there applies here too: engineered marks are improbable under independent genesis by construction, which is what makes their inheritance evidential, whereas a naturally-occurring habit is not demonstrated to be inherited merely because it is shared. Trap streets and dictionary Mountweazels were deliberately inserted conjunctive errors, planted so a copier could be proven a copier [Agloe, NY — a fictitious hamlet placed by General Drafting Co. in the 1930s as an anagram of its founders' initials; "esquivalience," invented by Christine Lindberg for the New Oxford American Dictionary, 2001, and later found republished elsewhere]. The term of art matters here: a ghost word is an accidental lexicographic error — "dord," a misread annotation that sat in Merriam-Webster's Second New International from 1934 to 1947, caught nobody and proved nothing — whereas a Mountweazel is planted deliberately, which is the only version that functions as a copying-detector. Watermarking and antidistillation fingerprinting are the identical design — synthetic Leitfehler — and their demonstrated inheritance is the cleanest existing proof that the descent channel is real.
The security consequence. Code is part of the corpus. A defect idiom emitted by a model in
year N is committed to public repositories, scraped, and trained into year N+2 models. The
inheritance channel is demonstrated for planted marks — watermarks, improbable under
independent genesis by construction. Whether it extends to naturally-occurring characteristic
habits is exactly the question §4 leaves open, not a second demonstrated case; treat it as
unresolved on the current evidence, not as a second Leitfehler-grade example. Its further
extension to defect idioms is an inference from the same mechanism, stated as
such and testable by retrospective tracing — if it holds, vulnerability patterns can be
inherited through the archive, not merely regenerated; and the channel is targetable: seeding
the corpus with a subtly defective idiom poisons the exemplar a generation of models will copy
[inference from the demonstrated channel; no direct study yet — verify against any emerging
literature].
Code is also, and this matters for feasibility, the best-conserved textual domain in the
tradition: version control gives every serious repository a hash-chained, signed, timestamped
manuscript apparatus of its own. The pre-2022 state of a codebase is not a lost archetype but a
checkout; dating is a diff against a signed tag. If prospective conservation is winnable
anywhere, it is here — and prospective is the operative word: unlike every philologist in
history, who reconstructed from the far end of a millennium of loss, this tradition can be
instrumented from its origin, with hashes and sealed witnesses from the founding stratum. The defence is conservation — provenance, dating, collation
against the oldest witnesses. Any empirical study of this channel must trace, never seed:
deliberately planting defects in public corpora is an attack, not an experiment (ethics
statement, §13).
Caveats carried with the section: the synthetic fraction of current crawls is uncertain and contested; frontier labs filter aggressively; nothing here is an agency claim about models — transmission structure, not intention.
The time-indexed thesis follows. The reviewers of our previous draft were right about the ecology their objections described — the error channel, where attractors dominate today (§4), and the terminal-reader era whose corpus shaped every present model: against that ecology there is no tradition to reconstruct and the Latin is costume. The recursion is young; the inherited fraction is small now and growing, which is precisely what makes the thesis time-indexed rather than wrong-then-right. The recursion falsifies each objection forward: the archetype (uncontaminated content) becomes practically lost as slop displaces it; the stemma becomes buildable because descent becomes real; the inherited fraction of shared features grows along a proven channel. Stemmatics exists only because traditions corrupt. A tradition has begun to corrupt. That — not any identity claim — is the case for agentic stemmatics, and it is a case that strengthens with time, which is an unusual and falsifiable property for a framework to have (§11 states the falsifier).
The previous draft was justly criticised for performing no distinctively stemmatic operation. This section specifies them as executable work with available ground truth. None is claimed as done.
7.1 The stemma of the models. This experiment performs recensio -- tracing lineage from
extant outputs, not conjecturing a reading no witness preserves; what that excludes is stated at
§11.7. Reconstruct the descent tree of a model population from output
evidence alone — conjunctive idiosyncrasies, per §4's corrected rule — and validate against
known lineages: base-model → fine-tune families, teacher → distilled students, checkpoint
successions, where ground truth is a matter of record. On present instruments, per §4's split,
this experiment has exactly one clean character class today: planted-mark tracing plus
reference-based teacher tests against known candidate lineages (distillation detection, teacher
identification — Rawat et al., arXiv:2607.09692) — legitimately a fingerprint-phylogeny
benchmark on its own terms. The distinctively stemmatic operations this section was criticised
for never performing — collation across witnesses, conjunctive grouping of natural characters —
become operable only once per-item improbability for natural habits is itself measurable, which
it is not yet. The full reconstructed-and-validated stemma appears open
[verify: absence claim from a sampled search, scope declared]. Failure modes are
informative philology: multi-teacher distillation and synthetic-data admixture are contaminatio
formalised, and measuring where reconstruction breaks is part of the result. The experiment is
cheap relative to its evidential value, requires no weights access for the subject models'
outputs, and its negative result would materially damage this paper's thesis — as a test should.
7.2 The independence metric for review panels. Replace assumed vendor-independence with measured seat-independence: per-question-class error correlation and idiosyncrasy distance between panel seats, tracked over time. This addresses the reviewers' finding that our previous "response-envelope" identity check measures product identity, not error covariance — the envelope tells you who answered; only measured covariance tells you how independent the answers are. The metric is proposed, not built; its own error characteristics would need the same scrutiny as any instrument in §8, and the deeper Bédier limit (§11) is narrowed by measurement, not closed.
7.3 Conservation: the crypt and the degradation clause. Frozen model weights are sealed witnesses: a model trained by year N cannot inherit year-N+1 contamination. An archive of open-weight vintages (closed models cannot be archived by third parties — an asymmetry worth stating: the open-weight ecosystem is the archivable tradition) supports three instruments: temporal decorrelation for panels (a sealed vintage cannot inherit post-seal contamination — though convergent attractors shared before the seal remain shared, so the decorrelation is on the inheritance channel only, per §4); normalisation-drift probes (an idiom that entered the tradition in year N reads as normal to year-N+1 models but remains high-perplexity to earlier, sealed witnesses — a dating instrument); and drift baselines that make corpus-degradation claims measurable rather than atmospheric. For operating panels we adopt a four-point degradation clause: (1) temporal stratification of seats; (2) the independence metric of §7.2, monitored, never assumed; (3) an oracle-share ratchet — the fraction of review weight resting on the mechanical oracles (execution, proofs, tests; the normative-text case of §3 is deliberately excluded here, since a ratchet toward document-reading is not the migration this clause intends) may only increase; (4) out-of-tradition witnesses (human specialists, formal methods) mandatory at the highest tiers. A method that measures its own decay can manage it; only the unmeasured version degrades silently.
Our operational practice maintains a catalogue of recurring instrument defects — ways a verification tool reports a false result, predominantly false absence: quiet success read as absence; matches discarded by the reader's own filter; summary output read as hits; a control exercising a different command than the claim; control class differing from corpus class (an open family — the skip rules of an arbitrary search cannot be enumerated in advance); patterns encoding a form the corpus does not use; shell/runtime dialect divergence. Each entry carries reproduced instances. Two are worth reporting here for what they evidence.
An absence-warranting gate that failed by its own catalogue. We built a tool to refuse "not found" claims lacking a known-positive control. An independent reviewer constructed six routes to a false "absent," including an inverted warrant (a quiet-success command reported absence precisely when the target was present); the tool's own five-case self-test had passed over all six, one control having been tuned — by its author — to a string that could not trigger the defect. The tool is quarantined, its defects encoded as permanently failing tests. We report this as a documented instance — one instrument, fully autopsied — of the hazard that verification instruments can inherit the disease they check for; the design rules it motivates (controls must be able to fail; graders graded against data their authors did not construct) are engineering responses to a demonstrated possibility, not conclusions from a measured rate.
A live recurrence during this paper's preparation. The operator asserted a document "does not exist" on the strength of a local-directory search while the document stood on the public web — control class ≠ corpus class, the catalogue's own fifth entry, committed by the catalogue's author in the closing hours of drafting. Absence claims in this paper are meant to carry their searched scope explicitly; where one below does not, that is itself an instance of the catalogue's own failure mode, not an exemption from it, and is corrected on discovery rather than defended.
The arithmetic of the one artifact-backed number, corrected. An author-selected mutation set had reported "5/5 killed — 100%"; an AST-enumerated run on the same module yielded 281 mutants of which 143 survived (50.9%) — a kill rate of 49.1% (138/281). The previous draft mislabelled the surviving count with the kill percentage; two reviewers caught it independently. We retain the observation as a single, artifact-backed illustration of one claim only — that generator-selected metrics are the generator's opinion in executable form — and draw no wider generalisation from n=1.
Two dated instances, kept distinct rather than filed as one: the oracle settling what a
panel could not, and, separately, calibrated convergence executed rather than prescribed. A
critique package arrived against a sibling instrument of this practice (the pre-build protocol's
frozen experiment commitments, not this paper), asserting among other things that the instrument's
central novelty claim was false.
The claim was checkable: primary-source verification (a standards specification — SARIF; two
public issue trackers — sarif-spec#120 and in-toto/attestation#77) substantially settled it in
one pass, with edge corrections, where the critique's own three commissioned investigations had
only asserted it — a correlation we do not measure and therefore do not claim of them either.
This is the first instance, dated and filed rather than hypothetical — an application of the
discipline §5's conditional recommends (oracle over panel where an oracle exists), not a claim §5
itself states: an oracle question answered
by one citation — two different things, properly: the specification is normative-text
consultation, an oracle in §3's sense deciding what the text says and nothing about whether it is
right; the two issue trackers are filed public records, observed rather than consulted. Neither
settles a claim about the world. Together they outperform three agents of one commissioned run,
however convergent. That is the first instance. Second, and separately: the
critique itself was then put through a three-lineage external seat round — mutually blind,
byte-identical brief, this practice's own prior recommendation deliberately withheld so no seat
could inherit it — and a majority of the three lineages rejected that withheld recommendation on
textual grounds neither had seen the other reach; the third lineage landed partial, permitting
the recommendation only as an explicitly declared stretch of the governing clause
[internal record, cited not reproduced — the sibling instrument's own filed record]. The
verification pass itself carried four minor inaccuracies (an enum undercounted, a field's
optionality inverted, an unconfirmed attribution, an uncounted variant) — found by that same
verification pass before use, not concealed — so we say substantially settled, with edge
corrections, not settled; a paper about unwarranted claims should not overclaim its own
verification. We do not offer this as evidence for this paper's thesis; the practice remains
one self-observed commercial operation (limit 10) and the critique concerned a different
artifact. We offer it as a dated demonstration that the discipline this paper prescribes —
oracle over panel where an oracle exists, convergence weighted by whether it answered a
withheld or a stated question — was executed, cost real verification work, and reversed a
conclusion two prior internal passes had reached. That is what practising the method looks
like, separate from whether the method is right.
Internal observations previously offered with load-bearing intent (rerun label instability; same-family blind-spot anecdotes; capability-level claims about self-clearing) are hereby downgraded to what they are: unaudited motivating observations from a single commercial practice, with the selection, incentive, and reproducibility problems that provenance implies (§11). Where external literature now covers the same ground (§4, §5), the citations carry the weight.
The previous draft claimed "the experiment is specified." That was false when written. The present
status is materially different and precisely bounded, and we quote the commitment record rather
than paraphrase it [internal record: PBM-VALUES-V8, frozen 16 Aug 2026 after an eight-round
adversarial review — 30 findings, 29 closed by the raising seat, one frozen as a declared
residual, none by vote; version chain v7→v8 documented with named record files].
How the freeze itself was governed. The freeze followed a pre-committed stopping rule for the review loop: the commissioner declared, on the record and in advance, that round eight was final and the freeze would occur regardless of that round's outcome, any new findings entering as declared residuals rather than amendments — so the record was not iterated until its reviewer tired. The version chain (v7→v8, one amendment applied under the prior round's "amend-then-freeze" verdict) is documented in named record files.
What is frozen (Values 1–4 of the protocol's commitment annex), enumerated:
Discontinuation is bounded: at most one revision, the failure filed first, re-registration with a new pre-committed threshold; a second failure discontinues; results are filed either way. Amendment of frozen values is barred — an amendment attempt is itself a discontinuation-clause event. Two governance flags recorded against this clause by the protocol's own reviewers remain unadjudicated, and we surface rather than resolve them: whether Values 1–3 are protected at a failure-triggered re-registration is unstated in the texts; and the confirmed residual on the ledger-close event (a late-submission refill surface) has a recorded candidate fix held behind the amendment bar rather than silently applied.
Status update, 17 Aug 2026. The values above are discontinued, not failed — no registered run occurred under them. A critique arrived, was substantially settled with edge corrections (§8's phrase, carried here deliberately rather than upgraded), and was put through a three-lineage external seat round.
A conflation in our own record, corrected rather than defended. An earlier draft of this section used the 2-to-1 majority's agreement to settle what the bounded-revision clause means — whether it keys on a run failure or an instrument failure — and then labelled that use "governance" to exempt it from §5's condition. That was wrong on its own terms: the clause's meaning is a truth-candidate, a descriptive claim about a text that could in principle be checked against the clause's drafting history or contradicted by a fourth reading. A correlated three-lineage majority is not a safe way to settle it, and reading the reasons before counting the votes does not change that — two correlated lineages agreeing after reasoning is still a vote with extra steps, not a different epistemic act. We withdraw the implication that the vote established the clause's meaning; the interpretation stands as an open residual, the third lineage's reading retained rather than outvoted, and nothing here adjudicates between them.
What the decision actually rests on. Something narrower, which does not depend on which reading is right: the instrument's governing rule was contested — an event visible in the split itself, not a finding about the clause — and continuing to operate an instrument whose rule is contested cost more than its remaining value. The three lineages' correlation (shared artifact, shared commissioning frame, plausibly the training-corpus hyparchetype §12 names for this paper's own convergence) is a reason their agreement carries little weight as evidence about the clause; it is not a reason against discontinuing, since that decision never depended on their agreement being correct. The commissioner accepted the discontinuation on those grounds. We no longer offer a general account of when aggregation may be used this way — §5 records why that attempt was withdrawn — and this section is an illustration of the discipline being applied, not a demonstration that it works.
What a reader cannot check here. The disputed clause is described in this section but never
quoted verbatim, because the record it belongs to is internal and cited rather than reproduced —
the standing limitation §12 and the References both record. The ambiguity is at least visible in
the paraphrase given above, which conditions discontinuation on "the failure filed first" and on
"a second failure" without anywhere stating whether an accepted critique against the instrument's
design counts as a failure at all. But a reader wanting to adjudicate the interpretive dispute
independently cannot do so from this text, and should treat our characterisation of it as
unaudited. We name this rather than leave it for a reader to discover, in the one place where the
non-reproduction gap bears directly on a decision this paper reports as made. The values in this section
stand as a closed record of what was frozen and why it was retired, not as an active
commitment. A successor instrument, built to a narrower kill-screen standard — externally-validated
ground truth, a frozen mechanical matching rule, blind non-author scoring — is a separate
commissioning, not a revision of this one. The reviewing lineages converged on discontinuing the
original instrument; the successor's specific design is this practice's own response, not a design
the seats themselves converged on
[record: PBM-DISCONTINUATION-RECORD-2026-08-17.md; the successor's own registration record
is still owed — this pointer gains a second target when it exists].
What may not be claimed, per the record's own exclusions: a keep verdict licenses only "this
builder-plus-protocol bundle was a better purchase than this baseline under these rules" — not
"discipline works," not transfer to other builders, not superiority over unconstrained
generation (the baseline is rule-governed). The case study is frozen as non-decisional: it
prices, and decides nothing. The registered multi-pair run (N = 4 pairs minimum from a
to-be-frozen eight-module pool) does not yet exist: the record states verbatim that the pool
"does not exist yet; the skeleton binds its shape and size." The record carries a nine-item
open-residuals inventory, including the permanent lineage/discipline confound, the partial
blindness of the finder seats, and an expectancy leak from case study to registered run. The four
frozen values map onto the protocol annex's required commitments (1)–(4); whether they have been
separately lodged with the recorder as a registration act is [verify] — the paper claims the
freeze, not the lodging. The transfer assumption this experiment exists to test is itself a standing open
finding against the protocol (its D5): our strongest prior evidence is review-brief evidence,
and its transfer to generation is unverified.
This is, we submit, what a preregistration-forward posture looks like from inside a practice rather than a laboratory: values frozen and quotable, boundaries of licensed claim frozen with them, and the distance to a real registered run stated in the same breath — because a paper that argued fluent overclaiming is the failure mode, and then overclaimed its own experiment, would be its own counterexample. It was, once. §12 records what that cost.
AI Control. The control agenda develops protocols around untrusted models — trusted monitoring, trusted editing, untrusted monitoring [Greenblatt, Shlegeris, Sachan & Roger, "AI Control: Improving Safety Despite Intentional Subversion," ICML 2024, arXiv:2312.06942] — reaching "assurance structure around an untrusted generator" from an alignment starting point. We reach the same architecture from software assurance and philology. We previously called this "decorrelated corroboration" and withdraw the inference: both lines descend from the shared post-2022 untrusted-LLM discourse, and convergence among correlated intellectual lineages is weak evidence — by our own §4. The convergence is noted; nothing is claimed from it.
Formal-verification-gated generation (proof-checked synthesis; state-funded mathematical assurance programmes - the UK's ARIA "Safeguarded AI," a GBP 59M state programme pairing world models with machine-checked proofs to yield quantitative guarantees) is the oracle-space limit of this program, applicable where machine-checkable specifications exist. Our degradation clause's oracle-ratchet is a deliberate migration toward that limit.
The scanner-era instrument stratum. A family of pre-LLM attestation formats already
solves adjacent problems for deterministic reviewers: SARIF (OASIS Committee Specification
2019, OASIS Standard 2020) defines a six-value per-result kind enum — fail ("matched"),
pass ("clean"), and notApplicable among them — and an artifact-role vocabulary marking what
was scanned, with content hashes attached; Microsoft's BinSkim and the OSS Review Toolkit ship
comparable per-target disposition and exclusion-reason machinery in production
[BinSkim's --kind filter takes Fail/Pass/Review/Open/NotApplicable/Informational per scanned
binary; ORT's PathExcludeReason enum records why each excluded path was excluded - nine reason
codes including BUILD_TOOL_OF, TEST_OF, DOCUMENTATION_OF. They classify different things, a result
disposition and an exclusion rationale, and are comparable in form rather than identical]; FOSSology has carried per-file human
clearing decisions since its open-sourcing in December 2007. These are not this paper's competitors; they are its correct
ancestors, and the honest lineage runs through them rather than through any bespoke schema.
Their shared assumption is the one this paper's subject breaks: a scanner's intended coverage
is deducible from its configuration, so a separate totality proof looked unnecessary. That
inference is weaker than it sounds, and this paper's own §8 catalogue supplies the counterexamples
— runtime failure, unsupported formats, silent exclusions, early termination, and caching all
break the step from configured to examined. Configuration establishes intent, not completion. A reviewer whose coverage is not deducible from
anything — a human skimming, or a model lineage whose attention over a tree is not observable to
the consumer of its output, and whether it is observable to the model itself is the open
interpretability question §1 leaves open, not one foreclosed here — has no analogous instrument. The in-toto attestation framework's own
maintainers have asked for one publicly, in the framework's attestation repository: issue #77,
requesting a generalized human-review predicate, has been open since 2021-12-05, with
review-coverage chaining across a diff history named explicitly as the unresolved design
question. This
paper's collation apparatus — closure as a verification property rather than a progress bar, a
property no section of this paper yet specifies as an artifact; §7.2 proposes an independence
metric and §9 records a ledger-close rule belonging to a discontinued instrument, neither of
which is the verifier meant here — is positioned into that stratum, not against it: what this
paper contributes at present is the requirement the gap names, not an instantiated verifier; the
intended packaging, if built, would emit SARIF results and an in-toto predicate rather than
inventing new interchange, and would claim only the closure verifier and the judgment predicate
as new. No such proposal has been submitted;
disclosure and timing are separate decisions this paper does not make [verify: pointer to the
submission record, if and when it exists].
Multi-model review systems — consensus aggregators, convergence-loop reviewers, panel pipelines, named with their documented aggregation behaviour in §5 — are the nearest mechanism-level prior art and the direct addressee of §5's conditional. Mechanistic interpretability is the complementary science, and the relationship is best stated as a division of labour: interpretability is the physiology of the organism — weights-access required, jurisdiction ending at each lab's own API, coverage presently, by its own practitioners' account, a small fraction of the computation. We found no published figure quantifying that share — an absence claim from a sampled search of the interpretability literature and lab writeups, not an exhaustive one, and scoped here per §8's own requirement rather than asserted flat; the further claim sometimes made, that exhaustive feature extraction would cost more than training compute, we have seen stated by practitioners but not seen quantified, and do not rely on it; agentic stemmatics is the epidemiology of the textual population — outputs only, cross-vendor by construction, its raw material compounding as the tradition deepens (its instruments, per §7, proposed rather than proven). The asymmetries favouring the epidemiological program for the output- warrant question are: jurisdiction (no lab can open a competitor's vat; the tradition is only studiable as a tradition); scaling sign (the interpretability gap plausibly widens with model scale, while collation's raw material compounds); institutional admissibility (a regulator can re-run a collation; no outsider can re-run a feature analysis on closed weights); and historical sufficiency (philology reconstructed the classical canon with a purely behavioural error-model of the scribe — mechanism-knowledge was never a prerequisite for warrant-knowledge). The precedent is epidemiological: the Broad Street pump handle came off decades before the organism was identified [Snow removed the Broad Street pump handle in 1854 without knowing the organism; Koch isolated Vibrio cholerae in 1883 - twenty-nine years later]. Where interpretability is strictly stronger, stated fully because the comparison is propaganda without it: behind unanimous inherited corruption (the perfect hyparchetype leaves no dissenting witness — only mechanism or an oracle sees behind unanimity); on agent-intent questions (deception, sandbagging) that may leave no cross-witness trace; in causal intervention and training-time repair; and prospectively, in mechanism-grade absence claims, which sampling witnesses can never deliver. The unfaithful chain-of-thought result — the scribe's self-commentary diverging from his computation — is an interpretability finding this practice depends on (Turpin et al. 2023; Lanham et al. 2023); the flow runs both ways, and the division of labour has a direction: collation is designed to be the cheap, continuous, ecology-wide triage that tells the expensive microscope where to look — a division this program proposes and has not yet operated at scale. One further relationship is worth stating plainly: the collation records this program would accumulate — dated strata, provenance manifests, sealed vintages, per-question verdict archives — are candidate validation ground truth for future mechanism-level work, the role industrial practice has historically played for the sciences that later explained it - Pasteur's fermentation studies were commissioned out of the French wine and brewing industries' spoilage losses, and the microbiology came out of the practice rather than the reverse [Pasteur, Etudes sur le Vin, 1866; Etudes sur la Biere, 1876].
Stated as the reviewers of the previous draft forced them to be stated, plus those the corrections introduced.
Version 0.1 was reviewed by four seats of distinct vendor lineages under an identical adversarial brief, with the author's self-assessment appendix stripped from the review copies. All four returned MAJOR-REVISION; none returned SUBMITTABLE or REJECT. Findings were extracted mechanically from the four filed records into a 209-row inventory (findings, must-verify items, and incorporation requirements), archived alongside this draft; this version is the disposition of that inventory, and the raising seats — not the author — hold closure over their rows.
Calibration, applied to our own evidence per limit 6: convergence among the four on questions the brief asked (the identity overclaim, smuggled efficacy, desk-reject risk) is partly correlation-by-commissioning and is weighted as such. The stronger — not unweighted — signal is the brief-independent convergence: two seats independently found that no distinctively stemmatic operation occurred in the draft; two independently caught the same arithmetic mislabel; four independently converged on the adjudication-DoS critique of §5. Even these share the artifact, the genre of reviewer, and plausibly the training-corpus hyparchetype (§11.5–6), so "stronger" is a comparative within one commissioning, not a claim of measured independence. Those findings shaped this version's structure — §4, §7, and the §5 symmetry are their disposition.
We offer this section as evidence for exactly one claim: that the method is applied to its own artifacts, at cost, with records. It does not answer the closed-loop critique — the author commissioned the review, wrote the brief, and integrates the findings; the review reduces self-certification, it does not abolish it (limit 4). A sibling critique, reviewed by three external lineages mutually blind to each other's output, against a different artifact of this practice, named the mechanism precisely: a closure that can be re-prompted, re-contextualized, and re-run until it succeeds measures the author's persistence, not the finding's resolution, unless the record captures every attempt — not merely the one that closed — outside the author's sole control. That finding applies with undiminished force to the four-seat closure this section describes, in a specific and checkable way: the raw seats' final verdicts are filed and byte-identical to what closed the rows — a reader can inspect them — but nothing in the record distinguishes a first-pass MAJOR-REVISION from a fifth attempt at the same prompt, because attempt counts, discarded runs, and prompt deltas were never captured. The gap is in the history behind the artifact, not in the artifact itself. We name it rather than paper over it.
What was promised, what happened, what exists now. The standing-practice fix proposed against the sibling artifact — an append-only log whose entries the author cannot be the sole gatekeeper over — was said to be "adopted going forward for this paper's own revision rounds, starting with whatever round follows this one" at v0.3. Seven revisions followed. None were logged. That is not a delayed rollout; it is the exact persistence-critique this section describes, committed by this section against itself, and it is corrected here by being stated plainly rather than by a further promise. The historical record — every round from v0.1 through v1.1, including the four-seat convergence this section reports above — will never be covered by attempt-level capture. That gap is permanent. A minimal version of the mechanism — an append-only, hash-chained log of seat invocations — was started at v1.1, covering rounds from that point forward only; it does not yet have external custody, meaning no party other than the author holds or can verify a copy, which is the specific condition the original proposal named and which remains unmet. A log the author alone holds is stronger than no log — it is at minimum tamper-evident to a second party who later obtains a copy — but it is not the independent record the proposal described, and calling it that would be the overclaim this whole passage exists to refuse. The difference between a manifesto and a field report is not that a field report keeps every promise; it is that a field report says which ones it broke, for how long, and exactly what replaced them.
A Bitcoin anchor, and precisely what it does and does not establish. On 20 August 2026 the
hash of this log's first entry was committed to the Bitcoin blockchain in an OP_RETURN output,
transaction 542748138c9889e745c99fbccd268ec16cafc3e53ebb6fe30b1ee44cdd0145e0, confirmed in block
963282 at 09:23:01 UTC. The payload is the ASCII string CPRLOG1: followed by the SHA-256
entry_hash of entry seq 1. Any reader can fetch that transaction, decode the output, and
compare it against the log file's own first entry. What this establishes is that this specific
hash existed, and was published, no later than that block — so the author cannot retroactively
alter entry 1, or the round it records, without the mismatch being detectable by anyone who
checks. What it does not establish is external custody, which is the condition the original
proposal named and which remains unmet: no party other than the author holds the log's contents,
and a hash is not a copy — it proves what a record was, only to someone who already has the
record. Nor does it speak to completeness. The anchor covers one entry, and the
log holds four: the three added after it were appended on 20 August 2026, reconstructed from
filed artifacts rather than captured at the time of the rounds they record. Those entries are
marked as retroactive in the log itself, because a line written a day late does not carry the
guarantee a contemporaneous one does — the artifact hashes in them are real, the timestamps of
their capture are not. They are also unanchored: only entry 1 is on-chain. This is a
tamper-evidence improvement over one entry of a four-entry log, not the independent record the
commitment described, and the distinction is exactly the kind this paper elsewhere insists on.
Published, and what publication changed. Later the same day the log was made public at
github.com/clvstra/agentic-stemmatics-closure-log, together with a standalone chain verifier and
the anchor record. The branch forbids force-pushes and deletions, with the rule applied to
administrators including the author; that was confirmed by attempting a history rewrite and having
the server reject it, not by reading the setting back — a control nobody has tried to defeat is a
claim, not a control. Commits are signed. A reader can therefore clone the log, recompute the
chain, check entry 1 against the Bitcoin transaction, and read the commit history for evidence of
later editing, none of which requires asking us for anything. What it still is not: the repository
sits under the author's own account, the second holder is the practice's other principal, and
GitHub is itself a trusted party to the arrangement. None of those three is independent of this
practice. The gain is narrower than "external custody" and worth stating exactly — silent revision
now requires defeating a public append-only history rather than editing a private file, and the
difference is checkable by strangers. Whether that suffices is not ours to declare: the finding
was raised by an external seat, and under this section's own rule a row closes only by the seat
that raised it. It is recorded here as addressed, and left open pending that disposition.
We claim: a corrected inference rule for lineage in model populations (idiosyncrasies, not errors, under the improbability condition); a conditional aggregation result grounded in external measurement, with our own alternative's costs stated; a diachronic argument that the model ecology is becoming a textual tradition through measured channels, making the stemmatic frame progressively exact rather than presently identical; an executable validation experiment with ground truth (§7.1) and a falsifier (§11.14); and a conservation program whose value is contingent on that thesis.
We do not claim: that the method produces better code (frozen experiment, discontinued before any run — §9); that adjudication should replace voting outside its stated operating region; that any internal anecdote establishes a rate; or that this paper's survival of its own review process evidences its truth. A preproof becomes a proof only by surviving collation it did not commission.
Ethics. The corpus-inheritance channel must be studied by tracing, never seeding: planting
defective idioms in public corpora is an attack on the commons this program exists to conserve.
Client-derived observations remain confidential and are marked unaudited wherever used. The
author's commercial interest in the method's adoption is declared and is a standing reason to
distrust the unexternalised parts of this record — which is why the load-bearing sections are
built to rest on external literature, frozen quoted commitments, and specified experiments
others can run. Whether they fully succeed in doing so is for readers and the raising seats to
judge, not for this paragraph to certify. A citation pass has since been run and its results are
in the References below. Seven claims carry [verify], named rather than counted: philological
term definitions against standard handbooks (§2), scribal-error directionality (§3), the usus
scribendi candidate anchors named at v1.8 (§4), the defect-idiom inheritance inference (§6), the
absence claim from a sampled search (§7.1), the registration-lodging question (§9), and the
submission-record pointer (§10). This count has now been wrong in six consecutive revisions, for
four distinct and separately diagnosed reasons — an em-dash-phrased tag not beginning with the
literal string "[verify"; a genuine fix that removed one tag while an unrelated miscount added one
back by coincidence; a tag whose text wraps across a line break, which every line-based automated
search this session ran, including the one that produced the immediately preceding wrong count,
could not see; and, this time, a tag a revision genuinely added — its existence stated in that
revision's own header — that the named inventory simply never absorbed. Two
independent full-text reads caught the third of those; a decorrelated seat, not the author,
caught the fourth. The count is not re-asserted as
settled; it is stated as the product of the most careful method tried so far, which is itself the
paper's own point about what "checked" can honestly mean.
Entries carry caps on what may be assumed of them, and no positive verification mark. An earlier draft of this section tagged every entry the author had checked against a primary source. Two external reviewers, asked independently, both recommended removing that tag: a status field only its author can populate is indistinguishable to a reader from a false one, and placing it beside specific figures implied a content check the artifact does not substantiate. The exhibit was inside this very section — a FOSSology date carried the tag and still contradicted the body text by a year. The marks that remain are limits rather than credentials. [SEC] = established from secondary or reference sources only. [PP] = preprint, not peer-reviewed at time of writing. [REPORTED] = rests on press or vendor reporting this paper did not independently audit. An unmarked entry is an ordinary citation, and claims nothing about who checked it.
Kim, E., Garg, A., Peng, K., & Garg, N. (2025). Correlated Errors in Large Language Models. Proceedings of the 42nd ICML, PMLR 267:30038-30066. arXiv:2506.07962.
Kim, D. (2026). Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles. arXiv:2607.20768. [PP]
Sun, M., Yin, Y., Xu, Z., Kolter, J. Z., & Liu, Z. (2025). Idiosyncrasies in Large Language Models. ICML 2025. arXiv:2502.12150. — five-way attribution reported at 97.1%.
Pan, L., Liu, A., Huang, S., Lu, Y., Hu, X., Wen, L., King, I., & Yu, P. S. (2025). Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? ACL 2025 (Main), 2025.acl-long.648. arXiv:2502.11598. — cited for the inheritance premise it measures, not for its robustness conclusion, which argues inheritance is defeatable. (Misattributed to "Zhao, X." at v0.5-v1.0; corrected on independent GLM/direct-fetch verification, both against the ACL Anthology entry.)
Rawat, R., Chen, S., Anand, A., Duan, M., Rotsted, B., & Min, S. Reference-Based Distillation Detection in LLMs. arXiv:2607.09692. [PP] — reference-based membership inference; despite adjacency in earlier drafts of this paper, it concerns no watermarking.
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature 631, 755-759. doi:10.1038/s41586-024-07566-y. Earlier: The Curse of Recursion, arXiv:2305.17493 (2023).
Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. [PP] — the accumulation-vs-replacement result; not part of Shumailov et al.'s own claims.
Kobak, D., Gonzalez-Marquez, R., Horvat, E.-A., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances. doi:10.1126/sciadv.adt3813. arXiv:2406.07016.
Juzek, T. S., & Ward, Z. B. (2025). Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models. COLING 2025. arXiv:2412.11385 — tentatively implicates RLHF; leaves the causal mechanism open.
Spennemann, D. H. R. (2025). Delving into: the quantification of AI-generated content on the internet. arXiv:2504.08755. [PP]
Graphite / Originality.ai analysis of 65,000 Common Crawl URLs, 2020-2025, reported Oct 2025. [REPORTED] — detector-based; detector reliability is itself contested.
Meta, Llama 3.1 Model Card; Microsoft, Phi-3 Technical Report, arXiv:2404.14219; Anthropic, Claude 3 Model Card. — cited for disclosed synthetic-data use in training pipelines.
Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023. arXiv:2305.04388.
Lanham, T., et al. (2023). Measuring Faithfulness in Chain-of-Thought Reasoning. Anthropic.
Greenblatt, R., Shlegeris, B., Sachan, K., & Roger, F. (2024). AI Control: Improving Safety Despite Intentional Subversion. ICML 2024. arXiv:2312.06942. — source of trusted monitoring, trusted editing, and untrusted monitoring. The "untrusted advice" protocol is a later, separate fellowship writeup and is not part of this paper.
ARIA (UK Advanced Research and Invention Agency). Safeguarded AI programme.
OASIS. Static Analysis Results Interchange Format (SARIF) v2.1.0. Committee Specification 01, 23
July 2019; OASIS Standard, 27 March 2020. — result.kind is a six-value enum:
notApplicable, pass, fail, review, open, informational.
sarif-spec issue #120, "Identify files that were scanned" (opened 2018-03-10, resolved into the artifact-role design).
in-toto/attestation issue #77, "Defining a generalized predicate format for 'human reviews' of artifacts" (opened 2021-12-05 by adityasaky; open at time of writing).
Microsoft. BinSkim, docs/UserGuide.md.
OSS Review Toolkit, PathExcludeReason (model/src/main/kotlin/config).
FOSSology, per-file clearing decisions since its open-sourcing, December 2007.
Wilson, P. (2026). The Star Chamber: Multi-LLM Consensus for Code Quality. Mozilla.ai, 5 March 2026.
Cursor. "Building a Better Bugbot." Cursor Blog. Cited for the aggregation mechanism as documented in the source (eight-pass majority voting); a self-published resolution-rate figure elsewhere in that source is not cited here, was not audited, and does no work in this paper's argument.
Benchmarking and Studying the LLM-based Code Review. arXiv:2509.01494 (2025). [PP] — SWR-Bench; Multi-Agg and Self-Agg aggregation variants.
Maas, P. (1927). Textkritik. English: Textual Criticism, trans. B. Flower, Clarendon Press, 1958. [SEC] — source of Leitfehler.
Timpanaro, S. (1963). La genesi del metodo del Lachmann. English: The Genesis of Lachmann's Method, ed./trans. G. W. Most, University of Chicago Press, 2005. [SEC]
Bedier, J. (1928). On the manuscript tradition of the Lai de l'Ombre. [SEC] — 105 of 110 surveyed stemmata were two-branched.
Cerquiglini, B. (1989). Eloge de la variante. English: In Praise of the Variant, trans. B. Wing, Johns Hopkins University Press, 1999. [SEC]
Roelli, P. (ed.) (2020). Handbook of Stemmatology: History, Methodology, Digital Approaches. De Gruyter. [SEC]
Roos, T., & Heikkila, T. (2009). Evaluating methods for computer-assisted stemmatology using artificial benchmark data sets. [SEC]
Reynolds, L. D., & Wilson, N. G. (2013). Scribes and Scholars: A Guide to the Transmission of Greek and Latin Literature (4th ed.). Oxford University Press. Orig. 1968. [SEC] — canonical general account of scribal transmission; cited for that, not as a confirmed primary source for usus scribendi specifically. (Duxfield's 2019 digital-humanities project page, cited in earlier drafts, uses the term correctly but is not the primary source; replaced on review.)
Pasquali, G. (1934). Storia della tradizione e critica del testo. Le Monnier. [SEC] — named by a philological review of this draft as the more precise candidate anchor for usus scribendi as a technical concept; not yet confirmed against the primary text.
Bischoff, B. (1990). Latin Palaeography: Antiquity and the Middle Ages, trans. D. Ó Cróinín & D. Ganz. Cambridge University Press. Orig. German 1979, 2nd ed. 1986. [SEC] — named by the same review as an alternative candidate anchor for scribal-habit terminology; not yet confirmed.
Snow, J. (1854), Broad Street cholera investigation; Koch, R. (1883), isolation of Vibrio cholerae. [SEC]
Pasteur, L. (1866). Etudes sur le Vin; (1876) Etudes sur la Biere. [SEC]
McDonald, K. (5 December 2022), first public use of the low-background-steel analogy for pre-AI-contamination data; Graham-Cumming, J., lowbackgroundsteel.ai. [SEC]
Agloe, New York — trap street placed by General Drafting Co. [SEC] — "esquivalience," coined by C. Lindberg for the New Oxford American Dictionary (2001) as a copyright trap. [SEC] — contrast "dord," an accidental ghost word in Merriam-Webster's Second New International (1934-1947). [SEC]
The following claims in this draft remain uncited and are marked in text: the paper's own internal records (§8, §9, §12), which are cited but not reproduced; the extension of watermark inheritance to defect idioms (§6), explicitly an inference with no direct study; the §7.1 absence claim about a full reconstructed-and-validated model stemma, which rests on a sampled search with declared scope; and the registration-lodging question in §9. These are named here rather than left for a reader to discover.
A note on this build. The drafting copy of this paper carries a working-notes section at the end — an internal maintenance list of open review rows, owed work, and venue thinking, marked in the source as delete-before-submission. It is deliberately not published here. Several entries below refer to it as "the Notes"; those references point at that omitted material. Nothing in the argument, the References, or this history depends on it, and the open findings it tracked are also stated in the body where they bear on a claim. The entries below are reproduced unedited, which is why each carries the "NOT SUBMITTED" stamp it had when it was written: they are dated records of what was true at the time, not statements about this preprint.
Draft v1.16 — 20 Aug 2026. Records the Zenodo deposit: concept DOI 10.5281/zenodo.22030516,
version DOI 10.5281/zenodo.22030517 for the v1.15 file, CC BY 4.0, published 20 Aug 2026. The
deposited PDF is Agentic-Stemmatics-Soons-2026-v1.15.pdf, 678,933 bytes, SHA-256
d70778bad4a6a6ed873f4db148f65bd053fa65b976e17217685a4f262dafcbd8. Zenodo freezes files on
publication, so that byte sequence is now fixed and independently checkable against the record.
The paper was not posted to arXiv: arXiv requires an endorsement from an existing author in the
subject class for submitters without an institutional affiliation, which this practice does not
have. That is a gate on visibility, not on priority, and it is recorded here rather than left as
an unexplained absence.
Draft v1.15 — 20 Aug 2026 — NOT SUBMITTED. Adds the author's ORCID (0009-0000-3088-6373) to the byline, so the identifier travels with the PDF rather than living only in a hosting platform's metadata. This revision's note was written directly into this appendix rather than at the top of the paper — the practice v1.14 adopted after the changelog had twice re-accumulated in front of the Abstract.
Draft v1.14 — 20 Aug 2026 — NOT SUBMITTED. This revision moves the notes for v1.11, v1.12 and v1.13 into Appendix A, where the rest of the revision history already lives. They had re-accumulated at the top over three revisions, putting a reader's first page of this paper on its changelog rather than its argument — the same drift v1.11 corrected once already. Nothing in those notes was reworded or cut; only their position changed. Appendix A now carries the complete provenance record from v0.1 through v1.13.
Draft v1.13 — 20 Aug 2026 — NOT SUBMITTED. The closure log is now public
(github.com/clvstra/agentic-stemmatics-closure-log), with force-pushes and deletions barred for
administrators including the author — verified by attempting a rewrite and having it rejected,
rather than by trusting the configuration — and with a standalone verifier anyone can run. This
addresses the external-custody finding open since v0.4 without closing it: the repository is the
author's own account, the second holder is the practice's other principal, and GitHub is a trusted
party, so none of the three is independent. Per this paper's own rule the row closes only by the
seat that raised it, so it is recorded as addressed and left open. §12 states the gain precisely
and refuses the larger claim.
Draft v1.12 — 20 Aug 2026 — NOT SUBMITTED. The closure log's first entry was anchored to the
Bitcoin blockchain this revision (txid 54274813...0145e0, block 963282, 20 Aug 2026 09:23:01
UTC), and §12 now records it together with an explicit account of what it does not buy. It does
not close the standing external-custody finding, which has been open since v0.4 and is named in
§12 and the Notes as still open: a hash proves what a record was, but only to a reader who already
holds the record, so no external party has custody of anything. The log had also fallen out of
use since v1.1; the two Grok closure rounds reported above were appended this revision, but
retroactively — reconstructed from filed artifacts, marked as such in the log, and not anchored.
A retroactively written entry is weaker evidence than a contemporaneous one, and the log now says
so on its own face rather than presenting four entries as though they were all captured the same
way.
Draft v1.11 — 19 Aug 2026 — NOT SUBMITTED. This revision moves the full revision history — previously the first thing a reader hit, before the Abstract — to Appendix A, after the References. Nothing in the history was reworded or cut in the move; only its position changed. See Appendix A for the complete provenance record from v0.1 through v1.10.
Draft v1.10 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. Grok's closure round on v1.9 returned
MINOR-REVISION (GROK-PAPER-V19-CLOSURE-2026-08-19.txt, filed by the operator from Grok's pasted
return). All seven v1.7 leftovers closed except F-v03-6 (the closure-log external-custody anchor),
which the round explicitly confirmed cannot be paid by wording and should stay open — no change
made to it. Two new findings on material no prior round had reviewed as live text, both fixed
here: (NF-v19-1) the v1.8 usus scribendi* [verify] tag was never folded into §13's named
inventory, which kept saying six when the live body carried seven — a sixth consecutive wrong
count, caught this time by a decorrelated seat rather than by either of this session's own
full-text reads, now corrected to seven with the omission's cause stated; (NF-v19-2) the v1.8
header claimed to apply "the philology-literate human read... owed since v0.4," but the questions
document that read was written for states plainly it needs a classicist and that no automated seat
can supply it — and the source of the answers this session worked from was never confirmed before
that claim was written. Asked directly this session; not yet answered. The v1.8 header is
corrected in place (original text kept, correction stated above it) rather than silently rewritten,
and the Notes item stays open until the source is confirmed. Grok also caught a citation hygiene
error in the v1.9 header (this note, before this correction, cited a wrong date on its own source
file) — noted here as the kind of small self-account slip this paper keeps finding in itself.
Draft v1.9 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. Grok's closure round on v1.7 returned
MINOR-REVISION (GROK-PAPER-V17-CLOSURE-2026-08-19.txt), closing two Mediums it had left open
across three prior rounds (F-v03-4, F-v03-5) and confirming NF-v06-1/2/5 fixed, but finding the
v1.7 subtractive cut honestly stated and incompletely executed: three new Lows (NF-v17-1/2/3)
where the header still narrated the withdrawn theory in the present tense, §8 pointed at a "§5
stopping line" the cut had removed, and the surviving §5 rule's trigger ("pooling correlated
judgments") was narrower than the conditional it was meant to serve ("unmeasured or failing"),
which would have let an unmeasured-but-uncorrelated pool through unchecked. All three are fixed
here, along with three Lows left open across multiple prior rounds: F-v03-2 (the "§5 asymmetry"
label mis-cited §5 and glued two separate dated instances into one — now separated, with the
mis-cite removed), F-v03-7 (§9 overstated the PBM reviewing seats as having converged on a
successor's design, when they converged only on discontinuing the original instrument — the
successor's design is this practice's own), and NF-v06-4 (§10's "unobservable even to itself"
closed a question §1 deliberately left open about the model's own internal state, not just the
consumer's — a phrase that survived one prior round undetected because it wraps across a line
break, invisible to a naive search, which Grok independently confirmed before fixing it). Not
fixed, and not fixable by wording:* F-v03-6, the closure log's external-custody anchor — naming a
party other than the author to hold or verify a copy is an operational decision this revision
cannot make for the practice; it remains open in the Notes for the author, honestly, rather than
closed by asserting an anchor that does not exist.
Draft v1.8 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. Correction, added at v1.10: the paragraph
below originally said this revision "applies the philology-literate human read... owed since v0.4."
That overclaimed. PHILOLOGY-REVIEW-QUESTIONS-2026-08-18.md states plainly it is written for a
classicist, textual critic, or codicologist, and that no automated seat can supply the read it
asks for; the source of the answers this revision worked from was never confirmed before that
sentence was written, and remains unconfirmed as of v1.10. The correct description is: this
revision used the questions document as a checklist and applied the ten answers it was given
against the live text, verifying what could be independently checked (bibliographic facts for two
candidate citations) and leaving what could not be (whether those sources actually treat the term
in question) as open [verify] tags. The owed classicist read itself — sourced, credentialed,
attributable — remains open. What follows is the original v1.8 note, otherwise unchanged.
Draft v1.8 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. This revision applies the philology-literate
human read of §2–§4 that was owed since v0.4 (PHILOLOGY-REVIEW-QUESTIONS-2026-08-18.md), answered
against the questions that document put to the reader. Of ten items plus a central question, six
were already correctly implemented by v1.7 and needed no change (the central question; items 1, 5,
6, 8, 10) — confirmed against the actual text, not assumed from the reviewer's own "Implemented"
label. Four required real edits, applied here: (2) a disclaimer that Maas's own Leitfehler
criterion was qualitative (unwahrscheinlich), not the P-notation this paper formalises it with;
(3) a sentence locating the homoplasy/synapomorphy apparatus correctly — the logical distinction at
the level classical stemmatics already required, the character-state machinery belonging only to
the discipline's modern computational wing; (4) an explicit "structural analogy, not direct
transposition" flag on the Bédier comparison, in both §2 and limit 5, since Bédier's finding was
about stemma-counting specifically and the extension to panel independence is this paper's own;
(9) the usus scribendi citation, which asserted Reynolds & Wilson as "the canonical account of
the concept" without that specific claim ever having been confirmed — softened to cite Reynolds &
Wilson for what is confirmed (a canonical general account of scribal transmission), with Pasquali's
Storia della tradizione e critica del testo (Le Monnier, 1934) and Bischoff's Latin
Palaeography (Cambridge UP, 1990) named as the reviewer's more precise candidate anchors for the
term itself — both bibliographically verified this revision (author, title, publisher, date),
neither yet confirmed to actually treat usus scribendi by name, so both carry [verify] rather
than being cited as settled. Item 7 (a mechanical sweep of §§2–4 and §6 for "reconstruction"
language doing work only emendatio* could license) was performed this revision: every instance
was checked against §7.1's and §11.7's existing recensio/emendatio distinction, and none overreaches
it — the sweep is closed clean, with no textual change required.
Draft v1.7 — 19 Aug 2026 — NOT SUBMITTED. UNGATED. This revision subtracts rather than adds. Across v1.3, v1.4 and v1.6 this paper attempted a general criterion separating decisions bound by §5's independence condition from operational choices exempt from it; each attempt was repaired by external review, and the third still had a hole. One reviewing seat noted that if the boundary needed a fourth structural iteration, the need itself was the finding. It did. §5's general theory and §9's application of it are cut — roughly eight thousand characters removed, the first net reduction in this document's history — and replaced by a stated withdrawal plus the one operational rule that survived: where a decision's justification depends on a factual premise settled by pooling correlated judgments, the condition binds that premise whatever the decision is called; where no such premise is load-bearing, it does not. §9 keeps the admission (the vote settled nothing about the clause; the discontinuation rests on the rule having been contested, an event rather than a finding) and drops the theory that was defending it. Nothing in §4, §6, or §7 — the paper's actual contributions — is touched. The prompt for this was a question about whether the review rounds had begun circling: they had, in one section, and the honest answer was to stop building there rather than to iterate a fourth time. v1.4 went to two seats independently. DeepSeek closed row I — open since v0.7 and the one this paper repeatedly called the one that mattered — judging the recursive decomposition to have fixed the displacement defect, and closed rows E and F as well. GLM verified the §5 literature live, the one check no text-only seat could run, and returned the sharpest finding of the round: the boundary was resting on the wrong citation. Judgment aggregation, in the formal literature, is defined by the objects aggregated (sets of yes/no judgments over connected propositions) and not by truth-tracking; its impossibility results concern consistency, not voter independence. Our appeal to it named a different axis than the one we needed. Confirmed independently against the primary source before acting, along with one supporting detail of GLM's own that did not hold — its claim that the entry's examples include normative propositions; they are factual and procedural. §5 is rewritten accordingly: the boundary is the older descriptive/prescriptive one, the independence requirement rides on Condorcet's Jury Theorem and the dependent-vote results (with two precisions GLM's check forced — sufficiently low correlation* rather than strict independence, and binary votes rather than pooled estimates), and the judgment-aggregation appeal is withdrawn as a misuse rather than quietly dropped. §5 also now concedes what it is not inventing: Popper demarcates theories rather than sorting judgments, the descriptive/prescriptive line is Hume's, and the recursion is Quine–Duhem's structure with the stopping line in the classical observation-statement role.
Both seats independently reached the same two fixes, which is convergence worth recording: the
stopping line must test a premise's content rather than its mode of access (a record of a
pooled judgment makes the pooling observable, never the conclusion observational), and §9's claim
that correlation strengthened the non-convergence premise had to go. That claim was ours and we
liked it; both seats independently identified it as having the every-outcome-confirms shape §11.13
already catalogues — correlated agreement discounted, correlated disagreement upgraded, no model
making either direction wrong — and GLM added the arithmetic, that at three seats and one draw
each a 2-to-1 split occurs around 38% of the time even on a strongly-peaked reading. Withdrawn to
the narrow claim correlation does not weaken the observation. Row B is fixed by broadening §3's
oracle definition to include normative-text consultation (both seats chose this over narrowing
§8), with the three carry-throughs GLM identified and we had missed: §7.3's ratchet, §11.2's
enumeration, and §9's "grading oracle," which was a third sense of the term hiding one section
away. §9's own premises are reclassified in the same pass. Rows B (pending confirmation), D, G, J
remain; J is now honestly bounded rather than argued, and §9's self-application is demoted from
demonstration to illustration because the disputed clause cannot be reproduced. v1.3 went to
DeepSeek, which returned MAJOR
CONCERNS (narrower than v1.1) and found a real hole in the boundary v1.3 had just inserted to fix
an earlier hole. The §5 test as first written — does the judgment forbid some state of the
evidence, or only prescribe an action? — scrutinised the speech act and not the premises
underneath it, so a governance decision resting on a pooled, correlated factual premise would
pass at the top while laundering a truth-claim one level down. That is the same defect the
boundary was written to close, displaced rather than removed, and the criterion did not survive
contact. §5 then stated the test recursively, with a counterfactual criterion for which premises
are load-bearing and an explicit stopping point so it would not collapse into "everything is
truth-reconstruction" — this recursive apparatus is exactly what v1.7 later withdrew (see the note
above); it is described here in the past tense because it no longer exists in the live text. §9
then applied the decomposition to its own discontinuation and, at the time, was judged to pass it
on the merits — notably because its second premise used the panel's disagreement as directly
observable data rather than any seat's conclusion as authority, a "correlation strengthens rather
than undermines" framing that DeepSeek and GLM later independently flagged as an every-outcome-
confirms claim (see below) and that does not appear in live §9. DeepSeek separately confirmed, by
its own counterfactual test, that §9's
cost/risk justification is genuinely independent of the withdrawn clause-interpretation claim.
Two of its other rows are also addressed: the disputed clause is still not quotable (internal
record, cited not reproduced) but §9 now says so plainly and points out that the ambiguity is
visible in the paraphrase it does give, and §10's unscoped interpretability absence-claim is
scoped per §8's own rule. Rows B, D, F, G remain open and are not text-fixable. A third
independent count of the [verify] tags landed on six, matching. A consolidated review report (source
unattributed at drafting, later confirmed as the operator's own synthesis of a philology
questions brief, a governance briefing, and a Qwen-authored memorandum whose bibliography was
found substantially degraded and is not itself cited here) proposed sixteen findings; ten are
applied this revision after independent checking, not on the report's word. Verified directly:
Timpanaro's actual argument (Bentley, Wolf, pre-Lachmannian origins — confirmed against the
book's own chapter contents, not just its existence); the Bédier/Cerquiglini conflation (now
separated as opposed rather than continuous reactions); usus scribendi's canonical source
(Reynolds & Wilson, replacing a blog citation — chapter locator not claimed, since that detail
could not be confirmed). Applied on the strength of their own reasoning: the codex descriptus
table row renamed rather than stretched past its logical fit; the hyparchetype disclaimer moved
to first use; a closing methodological statement on when the philological frame is dropped, not
just worn. The consequential fix at the time was §5/§9: a boundary (judgment aggregation requiring
independence versus preference/risk aggregation that does not, then grounded in the formal
judgment-aggregation literature) replacing the unfalsifiable "governance decision" label
DeepSeek's row I correctly attacked, with §9 no longer defending its 2-to-1 majority by
relabelling it. That specific grounding was itself later found to be a misuse of the cited
literature (GLM; see the v1.5→v1.6 note below) and the whole boundary apparatus was withdrawn at
v1.7 (see the note above). What survives into the live text is narrower: §9 withdraws the claim
that the vote settled the clause's meaning, leaves that interpretation an open residual, and rests
the discontinuation on independent cost/risk grounds that never depended on the vote being right.
Provenance of v1.1→v1.2 (18 Aug 2026): §12 rewritten that revision to state plainly
that its own v0.3 closure-log commitment went unbuilt across seven revisions, that the historical
record will never be covered retroactively, and that a minimal version now exists — hash-chained,
started at v1.1, no external custody yet — logging the invocation of this very DeepSeek round as
its first entry. v1.0 corrected the stale Notes-for-author
section that a second unattributed review (no lineage, no envelope, no terminal marker; still
unidentified when asked a second time) had read as a live "hard block" — the COI/affiliation
metadata itself had been correct since v0.7. That same review returned eleven further findings,
ten applied. Two seated reviews of v1.0 then followed — Qwen (formalising its own v0.4 rows and
v0.7 assessment; verdict CLOSED on both) and GLM (its first full-text pass after two v0.4
verdicts it had itself disclosed as bundle artifacts, withdrawn this round; verdict
MINOR-REVISION). GLM's live primary-source verification caught a real attribution error — the
watermark-inheritance citation's first author is Pan, not Zhao, independently confirmed against
the ACL Anthology entry before correcting it — and a fifth consecutive wrong [verify]-tag count,
this time for a third, distinct reason: the tag lives inside a sentence that wraps across a line
break, invisible to every line-based automated search this session ran, including the one that
produced the immediately preceding wrong count. Two independent full-text reads, not automated
recounts, caught it; both had also independently found the correct number where four prior
grep-based passes had not. Also applied: GLM's finding that the restructured §5 exemplar list
named three systems but left the discard-singletons mechanism with no cross-vendor exemplar,
stated as a gap rather than smoothed over; a softened §8 absence-claim commitment, since four
instances did not carry the scope §8 promised of them; and, on GLM's own suggestion, a
correlation-weighting sentence ported from §12 into §9, pricing the PBM discontinuation majority
at the same epistemic grade §12 already applies to this paper's own four-seat convergence — this
does not resolve DeepSeek's row I, it prices it honestly rather than leaving it merely disclosed.
Earlier fixes carried forward: the §4/§5 minority-filtering contradiction; the §6
watermark-versus-natural-habit correction; the §7.1 scoping sentence; the §10 artifact-claim
correction; undefined DeepSeek row letters and "no response envelope" glossed on use; and a
Cursor reference repaired after an earlier cleanup script had broken it, with an unaudited figure
dropped rather
than left as an unused floating claim. Not applied: the review's own commentary and philology
notes, since its lineage is still unconfirmed. DeepSeek's v0.7 findings — labelled B, E, F, G, I, J in that round's own
numbering, not otherwise defined in this text: an oracle-citation framing question, an
overstrong interpretability claim, the unauthenticated review history, the append-only log's
completeness gap, the majority-configuration tension in §9, and the unquoted governing clause —
remain untouched; all six want a
record or measurement outside the author's sole control that redrafting text cannot supply. Provenance of v0.1: four separately-commissioned
model seats of distinct vendor lineages — "distinct lineage" is a commissioning fact, not a
measured independence claim (§11.6, §12) — (Grok, Codex — driven; GLM, Qwen — operator-carried,
no response envelope — no machine-readable record of which served model, token count, or run
identity produced the return, only the operator's own filed header saying so) returned a unanimous
MAJOR-REVISION; v0.2 was their disposition, machine-verified by nine checker agents against all
209 extracted rows.
Provenance of v0.2→v0.3 (17 Aug 2026): four bounded additions from the PBM critique arc — §8,
§10, §12 (new paragraphs) and §9 (status note) — following verification, a three-lineage external
seat round, and an informed audit. Full record: PBM-DISCONTINUATION-RECORD-2026-08-17.md,
PAPER-V03-STAGED-EDITS-PBM-ARC-2026-08-17.md.
Provenance of v0.3→v0.4 (same day): v0.3 went through its first closure round against six
seat-passes — Grok, Codex, DeepSeek (full/delta access) and GLM twice, once driven on a trimmed
bundle and once hand-carried with live web access, plus a full-text hand-carried Qwen pass.
Verdicts split cleanly by access: every seat given the full disposition text landed
MINOR-REVISION or better on the pre-existing argument (Grok, Qwen; Codex closed 13 of 22 old
rows outright); GLM's two MAJOR-REVISION verdicts are substantially an artifact of a
bundle design that could not supply disposition text, disclosed as such by GLM itself both
times (structural, not content). Six concrete, verified fixes applied this revision: (1) §8's
"internally-correlated"/"correlated internal passes" language dropped — asserting correlation
without measuring it, even of a critique's own investigations, is the exact discipline this
paper argues against; (2) §8's "all three lineages independently rejected" corrected to state
the actual 2-firm/1-partial split (Qwen's own verdict on the underlying question was PARTIAL,
not a rejection — a Grok fact-check against the filed seat record, not a stylistic softening);
(3) §8's citations named directly (SARIF, sarif-spec#120, in-toto/attestation#77) rather than
by class, and the four verification inaccuracies' provenance attributed to the verification
pass itself; (4) §10's in-toto citation names the repository in reader-facing prose, not only
in an internal [verify] tag that would not survive to publication; (5) §10's SARIF timeline
precised (2019 Committee Specification, 2020 OASIS Standard) and its kind enum given as six
values with three named tokens, resolving both the earlier trichotomy imprecision and a
citation-repo ambiguity a hand-carried GLM pass caught by live primary-source verification;
(6) §12's closure-log commitment gained a self-aware caveat — the anchor itself is unnamed and
a further instance of the same problem if chosen unilaterally by the author. Two items
explicitly NOT touched this revision, named rather than silently deferred: formal
affiliation/correspondence/COI metadata (Qwen's new Finding 27 — needs real values, not
placeholder text) and the remaining [verify] tags (a citation-research task at a different
scale, now underway — see the v0.4→v0.5 note immediately below for what has been resolved so
far).
Provenance of v0.4→v0.5 (18 Aug 2026): the v0.4 closure round returned six verdicts (Codex
MAJOR-REVISION; Grok, DeepSeek, Kimi, GLM, Qwen MINOR-REVISION), converging on a short list of
small, confirmed defects — a self-contradictory [verify]-tag count (header said 40, notes said
38; three independent counting methods, Grok/Kimi/GLM/Qwen, converged on 35, not yet reconciled
into the text below pending the fuller citation pass this note describes), an Abstract sentence
that kept the unmeasured-correlation language §8's fix had dropped elsewhere (fixed below), and
a §9 sentence that said "was verified" where §8 says "substantially settled, with edge
corrections... not settled" (fixed below, matched to §8's own qualifier). One finding was not
mechanical: Kimi, reading the paper cold with no framing inherited from the other five seats,
found that §4's corrected inference rule applied its own improbability-under-independent-genesis
condition to the error channel but not to the idiosyncrasy channel it replaces errors with — the
cited 97% attribution figure measures discriminability (can models be told apart), not
improbability (would an independent model develop the same idiosyncrasy by chance), and the
paper's own §6 discussion of RLHF-driven vocabulary convergence (e.g. "delve") is a real,
citable candidate case of exactly the convergent-attractor problem this channel was supposed to
avoid. §4 is revised below to apply the improbability condition per idiosyncrasy rather than to
the channel wholesale, splitting engineered markers (watermarks — improbable by construction)
from naturally-occurring idiosyncrasies (which need the same per-item scrutiny errors get, not
yet uniformly available). A parallel six-thread citation-verification pass checked roughly twenty
[verify]-tagged claims against primary sources; most held, several needed real correction (not
just sourcing) — including a watermark-inheritance citation that had conflated two unrelated
papers, one of which argues the opposite of what it was cited for, and a wrong technical term
("dictionary ghost words," which properly denotes an accidental lexicographic error, not the
deliberate copyright trap the paper meant — the correct term is "Mountweazel"). The full
bibliography and remaining tag resolution are a separate, still-in-progress pass; this revision
applies only the items resolved so far plus the three items above.