A model hands you sound code and a confident note saying it is safe. It can hand you broken code with a note that reads just as well. One process writes both; that is how these systems work, not one model's flaw. This introduction sets out how we review code when the note cannot tell you which is which.
Most attempts to solve this problem reach for the same fix: ask more than one model, and take the answer most of them agree on. It sounds sensible, and it is often wrong in a specific, measurable way. Models trained on similar data, with similar methods, tend to make the same mistakes for the same reasons. When that happens, three AI reviewers agreeing is not three independent confirmations; it is one mistake, echoed three times, that now looks like consensus. Worse: if one of those reviewers happens to be right and the other two share the same blind spot, a majority vote throws out the correct answer and keeps the popular one.
This is not a new problem. Textual scholars have worked on it for at least five centuries.
Long before computers, scholars faced almost exactly this problem trying to recover the original text of ancient books. No copy of Lucretius's poem survived in his own hand: only a scatter of hand-copied manuscripts, each with its own scribal errors, each a copy of a copy of a copy. Scholars built a discipline, called stemmatics, for recovering a trustworthy text from a collection of individually unreliable witnesses. Its central insight, refined over centuries of getting it wrong first, was this: where independent copies agree, the agreement is evidence; where they agree because they share a source, it counts once. Telling those two kinds of agreement apart, which requires knowing how the copies actually relate to each other and not just counting them, was the entire craft.
That is the discipline this practice borrows, and the reason we call the method Agentic Stemmatics. A model's output is treated the way a philologist treats a manuscript: a witness, not a verdict: useful, but only once you know what it might be corrupted by, and only once you know whether it is independent of the other witnesses you're comparing it to.
In practice, this changes three things about how a review gets done.
We do not assume independence. Two AI reviewers from different vendors are not automatically independent just because they have different names on the box: if the training data, the prompt, or the framing they were given overlaps, their agreement is worth less than it looks. We record what each reviewer was given, so overlap is visible, but we have not measured how independent their errors are.
Disagreement is treated as information, not noise. A single reviewer who disagrees with the rest is not automatically wrong, and is not automatically outvoted. It is investigated, the same way a scholar investigates the one manuscript that reads differently from all the others, because that manuscript is sometimes the only one that preserved the truth.
Findings are adjudicated, not averaged. Nobody's opinion gets diluted into a score. A real defect, once shown, blocks, regardless of how many reviewers missed it. And every finding is attached to a named reviewer's judgement, not a black-box aggregate nobody can question afterward.
None of this makes the underlying models more reliable, and we do not claim it does. What it produces instead is something narrower and, we think, more valuable: a documented, inspectable account of what was actually checked, by whom, and on what basis, the kind of record that holds up when someone with a reason to doubt it (an auditor, an insurer, a regulator, opposing counsel) actually goes looking.
We are just as direct about what we have not yet proven. We have not shown that a panel finds more real defects than one good reviewer. When we measured our full multi-model review against one inexpensive pass, on two codebases, it did not earn its cost. An earlier controlled experiment designed to test the claim was retired before it ran a single case: an external critique challenged the instrument, our own review process could not cleanly settle who was right, and we judged that continuing to operate under a live, unresolved objection cost more than the instrument's value. We would rather report that than paper over it.
This document is the short version. The full argument (the theory behind why AI-generated errors behave differently from the ones philologists originally solved for, the evidence for when panel review is and isn't safe, and an honest account of this method's own limits) is in the paper this introduces: Agentic Stemmatics: Collation, Provenance, and Conservation for an Emerging Machine Textual Tradition. It has not been peer-reviewed, and it says so on its own first page.
Field notes. Shorter, dated records of one specimen each, with method, counts and limits. Field Note 1. Consensus as oracle: TypeSafe's Jev benchmark as a type specimen (16 September 2026) opens the series: a vendor benchmark whose reference labels are the mean of two frontier models, read from the twenty per-case files the vendor publishes. The two labelers disagree on the final action in 8 of 19 cases; every case the vendor marks "all three miss the reference" is one where the labelers missed each other. Eight outside reviewer seats from four model families recomputed every count before publication. A plain-language version is on counterproof.io.
The field now has a home of its own. Agentic stemmatics is larger than this practice: a proposed research field (the philology of machine-generated text) with its own open problems, its map of the adjacent literatures, and its scope stated as carefully as its ambitions. CounterProof applies part of it and funds some of the research; the field belongs to whoever does the work. Its reference site is agenticstemmatics.org.
CounterProof is a small practice of Clavestra Capital Limited (Malta, C 113987) that reviews AI-generated code and AI-assisted review processes for organisations where a wrong answer is expensive. We disclose, before anything else, that we have a commercial interest in this method's adoption, the same discipline this document describes.
The paper is also deposited as a citable preprint under doi:10.5281/zenodo.22030516, CC BY 4.0.