audit · retractions

Retractions

Every claim this programme has made and then withdrawn, with what killed it. Nothing here is deleted from where it originally stood — each is struck through in place, so a reader arriving at the original text sees the correction rather than a clean document that never made the mistake.

This file exists because the site quotes a count, and a count nobody can audit is worse than no count.

Standing tally: 21 claims withdrawn, five experiments void, 1 finding established.

The last two numbers were both stale until 2026-08-04, and in opposite directions. The void count missed E-001c; the findings count still said zero after E-002c landed, so this file went on advertising a cleaner failure record than the programme had earned. A file whose whole purpose is that a count can be audited had two counts nobody was auditing. check_counts.py now reads the void number here as well as on the site; the findings number is not mechanically derivable — an experiment with a FINDINGS.md and no VOID.md is not the same thing as an established result — so it stays a judgement, and it stays visible here.


Axioms and definitions

#ClaimKilled byWhere
1A3 — no bounded message closes an arbitrary prior gapA 113-token lookup table reaching F* = 1. Restated for held-out probes under bounded cost.PRINCIPIA §4
2D_floor is "the irreducible divergence caused by the parties' own stochasticity", and belongs inside the definition of F*Two perfectly aligned stochastic agents have identical true distributions, so their true JSD is exactly zero. The floor is estimator bias, and a fidelity that changes when you sample more is not well-posed.definitions §3
3η = F*/C is "the quantity engineering should optimize"A ratio with a signed numerator is not an ordering. At F* = −1, a 100-token antinoophor scores −10.00 and an 800-token one −1.25, so the message that spends eight times as much to do the same damage ranks higher. Replaced by V_λ = F* − λC.definitions §4.3
4Falsification criterion 2 — "Φ ≈ 0 means the pathology does not exist"Bias and resolution are independent. A party predicting 0.70 on every probe and averaging 0.70 has Φ = 0 and no ability to say which probes it got wrong; it is maximally pathological and the criterion scored it as our refutation.PRINCIPIA §7

Laws

#ClaimKilled byWhere
5L2 headline formTautological as stated.laws.md
6L4 — fidelity is multiplicative along a chainIll-typed: it multiplied fractions of different prior gaps, and two antinoophors composed to a positive product. Measured: hops of −0.629 and −1.000 multiply to +0.629. Restated as L4a/L4b/L4c.laws.md
7L6 is "the field's first engineering prescription"I-PASS, deployed and outcome-measured since 2014 across nine programmes and 10 740 admissions.laws.md · prior-art §5
8L5 is "our sharpest conjecture"The human half is Carpenter et al. (2013) and Deslauriers et al. (2019). Status stays conjectured — a prior is not a test — but the framing was ours to lose.laws.md · prior-art §4

Novelty claims

#ClaimKilled byWhere
9Φ is "the part we have not found elsewhere", and is what "everyone had felt, nobody had weighed"Keysar & Henly (2002), Newton (1990), Chang et al. (2010), Endsley (2020). It was weighed in 1990, and their instruments are in places better than ours.PRINCIPIA §1, §5 · prior-art §1
10Knowledge distillation "measures success as task accuracy"Stanton et al. (2021) define fidelity separately from generalization and show accuracy does not imply it — E-001's construct failure, from a NeurIPS abstract, five years early.PRINCIPIA §2 · prior-art §3
11"Fidelity-versus-correctness was separated in ML first"Cronbach (1955) separated an accuracy score from an assumed-similarity score, each with its own decomposition; Edwards et al. (2006) measured team mental-model similarity and accuracy as two quantities and compared them as predictors. Both predate the ML work. The content survives; the primacy word does not.prior-art §3

Findings

A section that should have existed since 2026-08-02. The eleven above are claims about the theory and its novelty; these are claims about a measurement, killed by a better measurement. The first of them was struck correctly at its origin and indexed nowhere, so the standing count was one short for two days and no check could see it — check_retracted.py did not read experiment documents at all until 2026-08-04.

The second is here on the same rule and not because it is comparable in weight. It stood for one day, in an instrument note, and was killed by widening the very measurement that produced it. A file that records only the expensive mistakes would be a curated list rather than a ledger, and the cheap ones are the more common failure.

The fourth, #15, is of a kind this ledger had not held before. The three above withdraw a quantity; #15 withdraws the warrant — the arithmetic was performed correctly on a sample that does not have the structure the test assumes. It therefore reaches backwards into #14, whose own supporting evidence was one of the p-values it strikes, and it cannot be repaired by recomputation: a measure built from one-token variants of a shared prompt has no defensible number of independent rows, only a range that moves with where the clustering line is drawn. The lesson is not about Fisher's test. It is that A2 makes the probe measure the frame of reference, and nobody had asked how many distinct readings a frame of thirty-two probes contains.

#ClaimKilled byWhere
12"The parties' confidence is very nearly unresponsive to how much actually transferred" — E-002b's headline pull-quote, labelled post-hoc when madeE-002c committed the quantity before collecting and measured β = +0.1299, CI [+0.047, +0.223], which clears zero. Confidence does respond, at about an eighth of the rate calibration requires. The direction replicated and the magnitude did not. The unresponsive party is the sender alone, β = −0.02 with its interval spanning zero — a narrower claim than the one withdrawn, and a sharper one.E-002b §6 · E-002c §3
13"Zero of the 25 non-interaction probes diverge at 230 words" — stated in Problem 15 from six messages, and read as the property that only interaction probes survive saturationWidening the same measurement to twelve messages across all four cells found three: M02, M12, M29. A small-sample zero, killed within a day by the instrument that produced it. What survives is the proportion — 21 of 24 divergence events on 9 of 34 probes — and the divergence rate the specification is costed against, 0.194 against 0.204, which the widening left standing.Problem 15 · headroom-2x2.json
14"MERIDIAN-IX32 reaches 3.33 diverged probes per message and clears E-002c's outcome-variation gate" — the candidate probe measure built to repair Problem 15The admission gate behind that number had been applied once. Applied five times, admitting a probe only if it returns the key at margin ≥ 8 in every pass, it rejects four probes, and the measure on the 28 survivors gives 2.00 per message — below the gate. Worse for the repair than for the claim: all four rejected probes are among the nine that ever discriminated, and none of the twenty-three that never discriminated was rejected, p = 0.0035. The headroom was borrowed from probes that are not stable observables.MERIDIAN-IX32 · Problem 15
15Every Fisher exact p computed over MERIDIAN-IX32's probes — p = 0.0035 and p = 0.0138 for the anti-correlation between discriminating power and modal stability, and p = 0.0339 for the design rule that earned X17X32The measure is not 32 independent probes. X17 and X18 differ by one token — "Combined figure count is 70" against "71" — and carry opposite keys; so do X21/X22 (8 → 7) and X24/X25 (30 → 21). Single-link clustering of the prompts gives 9 clusters at similarity 0.80–0.85, the largest holding 11 probes, and all instability falls inside 2 of them. Every one of these tests treats near-duplicate rows as independent draws. There is no corrected number to substitute, which is the point: the cluster-level p ranges from 0.111 at threshold 0.80 to 0.006 at 0.95, so the result is a function of a clustering knob rather than of the data. What survives is at the draw level and is stronger — of 1 440 gpt-oss:120b sender draws, 19 are non-key and all 19 fall on R5-tagged probes, a label committed 2026-08-10 09:29:56, nine hours before the first IX32 sender pass and therefore prior to every instability datum.MERIDIAN-IX32 · Problem 14
16"Seven of the nine discriminating probes measured the reader, not the transfer" — and its corollary that MERIDIAN-IX32 measures model-independent transfer loss on 2 of 32 probes because the other seven are one reader's uncertaintyThree independent grounds. (a) The statistic was not like-for-like. "Six of the seven sit at gpt-oss's wobble points" compares gpt-oss's minimum over four sender passes against qwen's single pass; more passes means more chances to show a low margin. Counted one pass each it is 2 of 7. (b) No reader-specific term is needed. A model with one difficulty per probe and a single reader-ability gap — no reader×probe interaction whatever — fits at deviance 10.54 on 8 df, p = 0.229, with the gap at 2.20 logits. The strict-subset structure is what that model predicts anyway. (c) It is impact, not differential functioning. Dorans & Holland (1992): DIF requires comparing examinees "supposed to be comparable with respect to the attribute measured"; an unconditioned difference between groups of unequal ability is impact, and Simpson's paradox is the named hazard. The comparison was never conditioned on the 2.20-logit gap. The counts survive — qwen loses 2 of 32, gpt-oss 9, qwen's set a strict subset — but why is open, and was published as answered.results · prior-art §11
17E-002c's published per-rung fidelity column — +0.228 / +0.456 / +0.606 / +0.848, and the three briefs reported at exactly 1.0000The noise floor was computed on a different pair from the comparison it corrects. runner.py:222 takes the permutation floor between the sender and PRIOR; line 241 then divides every sender-versus-receiver fidelity by it. D_floor is estimator bias for the pair being compared, and these parties are not alike — the sender is a point mass on 33 of 33 probes, PRIOR on 4 — so the mismatched floor is 0.0417 against matched floors of 0.0076–0.0226. Every fidelity was inflated, and because the matched floor falls as cost rises the inflation steepened the ladder as well as lifting it: corrected, the column is +0.217 / +0.424 / +0.560 / +0.784, the climb is +0.57 not +0.62, and the three briefs at the min(1.0, ·) cap become 0.912 / 0.915 / 0.953no brief on the ladder reaches 1.0. β, the registered primary quantity, is a slope of claim on observed agreement and never reads this column; §3's finding is untouched.DEFECT-001 · E-002c §4
18"The fluent register's length floor sits above the band's ceiling … the manipulation is unsatisfiable rather than underpowered" — E-001c's stated void reason, and the claim that "the floor belongs to the fluency axis"A second model satisfies it. The same script, specification, calibrated instruction and band, changing only the model: qwen3.5:35b puts cell A in band 11 of 12 against gpt-oss:120b's 0 of 12, and cell B 12 of 12 against 5. The fluent floor is 223 for gpt-oss and 197 for qwen — above the 231 ceiling for one, 34 words below it for the other. That sentence is itself withdrawn 2026-09-08, retraction 19 below: 223 is eight words below 231, not above it. The in-band counts either side of it are measured and stand. The fluency-axis claim goes with it: fluent cells land in band 5 of 24 on gpt-oss and 23 of 24 on qwen, against terse's 23 of 24 on both — so on qwen there is no fluency-axis length effect at all. The void itself stands: E-001c ran on gpt-oss, that model could not compose inside its own band, and the experiment correctly died. What is withdrawn is the generalisation from one generator to the register. L5/L6 are not revived — this measures the band filter only, and E-001c's gate was the band and two blind raters, whose intersection is where the 0 came from.RESULTS-qwen-floor · E-001c VOID
19"gpt-oss's fluent floor sits above [the band ceiling]" — the gloss under retraction 18's own evidence table, published in the results file and repeated in row 18 aboveThe table it sits under refutes it. RESULTS-qwen-floor.md reports gpt-oss's fluent floor, minimum over cells A and B, as 223. The ceiling is 231. 223 is eight words below it. The claim is true only of cell A (floor 232, 0/12 in band); cell B's floor is 223 and it lands in band 5 of 12 — so gpt-oss reaches the band under a fluency instruction, in the contrastive cell, and what it cannot do is fluent × declarative specifically, which is the cell E-001c's VOID names in its gate. That is a statement about length and not about register: this run measures the band alone, and E-001c's VOID records that of the two live cell-A messages that did reach the band, neither was judged fluent prose. qwen's "34 words below" is correct. Retraction 18 is untouched: it rests on the in-band counts — 11 of 12 against 0 of 12 in cell A, 12 of 12 against 5 of 12 in cell B — not on this sentence. A wrong gloss was carried on top of a right finding for six days, into the ledger and onto the front page, because the number under it was never subtracted from the number beside it.RESULTS-qwen-floor
20"No claude-* model has ever read RIVERSIDE-30" and "It has never been run on RIVERSIDE-30" — the founding premise of the 2026-09-01 crossover prediction, and with it the claim that the measurement was blocked on an absent API keyIt had been run on 2026-08-03, and the file that proves it is the file the same paragraph cites. E-004 measured claude-opus-4-8 against RIVERSIDE-30@2e6afe2f3c92 — the identical measure hash — at 0.733, beside gpt-oss's 0.967, in E-004-20260803T161541Z.json, committed 801ac86. E-001c's PARAMETERS had quoted that very figure since 2026-08-03. The prediction is scored and both its sharp clauses hold — below 1.000, and acc(claude) − acc(gpt-oss) negative on both domains (−0.233, −0.091) — so no crossover now rests on four models rather than three, and rested on enough a month before it was argued. What is withdrawn is the reason the question was thought open: the 2026-09-01 → 09-06 arc, 42 GB of weights and 1 280 probe calls, was launched to settle something already settled and filed. The exposure claim in the same file survives — a stateless subject is not a contaminated agent.RESULTS-crossover · E-004 VOID
21"it fails E-004's 0.90 subject gate on both [domains]" — published in theory/laws.md, on the front page, and in three probes/ documentsE-004 registered no such gate. Its pre-registration §5.1 sets each model's accuracy > 0.60 on each measure, and its own void note records that "every model cleared the 0.60 accuracy floor". On E-004's real gate llama3.3:70b passes MERIDIAN-34 at 0.824 and fails only RIVERSIDE-30 — so a reader who audited the citation found the opposite of what the sentence claimed, on one of the two domains. The 0.90 threshold is real and is the sender accuracy vs key > 0.90 gate registered in E-001b, E-001c, E-002b and E-002c; llama3.3 fails that on both. The conclusion survives, its authority did not: the same shape as retraction 16 in the Findings table above, where the counts held and the reason given for them did not.RESULTS-llama-crossover · laws L2

Void experiments

idwhat it was going to measurewhy it is void
E-001fluency vs contrastiveness on fidelity and ΦSender refused to compose. Reopened as a construct critique: the headline quantity rewards mimicry, and 62% of its effect sat on four probes where the sender was wrong.
E-001bthe same, factorially, with a cost-parity gateThe gate failed on the composed messages: fluent briefs cost 1.5× terse ones under an identical budget instruction. Style and length are entangled in the generator. Also DEFECT-001 — the analysis path had never been executed and would have crashed after 30 hours.
E-002Φ for the first time, elicited per probe330 of 330 elicitations returned "yes, we will agree" — the instrument has a default answer, because it ported Keysar & Henly's granularity and not their forced-choice structure. And the transfer was perfect (0 of 33 probes diverged), so there was nothing for a belief to be wrong about.
E-004whether a model can predict where another model will disagreeTwo of the three registered models were never wrong: 1 error and 0 errors across the whole design. A detector of disagreement needs disagreement to detect, and reporting the run would have reported the defect as the result.
E-001cthe E-001 question again, with length pinned by a two-sided word bandThe fluent register's length floor sits above the band's ceiling — 232 words against 231, at 0 of 12 and 0 of 40. Fluent cells land in the band 5 times in 24 against terse cells' 23 in 24. The fluency instruction is partly a length instruction, so the manipulation is unsatisfiable rather than underpoweredwithdrawn 2026-08-31, retraction 18 in the Findings table above: qwen3.5:35b satisfies the same band 11 of 12 in cell A where gpt-oss managed 0 of 12. The void stands; the generalisation does not. Also DEFECT-001 — the first of the two runs voided on a copied instruction the band was never calibrated on, and establishes nothing.
E-006whether F*(C) is concave — the ablation ladder for L1Void before collection, zero model calls. Its own pre-registered power gate, computed before any draw, returned 0.117 against a moderately concave truth on 32 probes. L1's concavity limb needs about 512 probes; MERIDIAN-IX32 has 32, of which about nine are independent prompt templates. More draws do not help and more rungs do not help — the noise is probe-level. What this establishes is better than the run would have been: L1 is untested not for want of trying but because no probe measure this programme owns can carry the test.

This table was two rows stale until 2026-08-04. E-004 voided on 2026-08-03 and E-001c on 2026-08-04, and neither was indexed here. The count on the front page did not notice, because it counts VOID.md files rather than rows in this table — so the number was right while the list a reader actually reads was wrong. check_counts.py compares numbers to their sources and has nothing to say about an index that omits an entry.


Refused entry

Drafted, checked, and never committed. They are not withdrawals — they never reached theory/ — but they are recorded so that nobody, including their author, drafts them again.

claimwhy it was refused
WOA attributed to Yaniv & Kleinberger (2000)That paper publishes the complement, `WOE = \a−f\/\a−i\`. Every number in it is a WOE; the citation would have inverted all of them.
"Normalized gain is known to correlate with pretest score"Contradicted by the paper cited for it. Hake (1998) reports r = +0.02 across 62 courses, and that near-zero correlation is his central justification for the measure.
The judge–advisor weighting/accuracy separation as "a fourth independent arrival at the fidelity-versus-correctness split"The field reports the two quantities separately but does not theorise a split. That reading is ours and may be recorded only as ours.
Edwards et al. (2006) cited for "accuracy predicted performance where similarity did not"Unsupported by the only text available. The abstract says accuracy was the stronger predictor, which presupposes similarity predicted too.
Problem 16, "a probe measure read by one agent confounds the frame with the reader"Drafted 2026-08-19 and refused the same day. It is generalizability theory's fixed facet, and Brennan (2003) states as a theorem that with all facets fixed "no generalization is involved, and all error variances are zero" — a known design cost with a known remedy (a D-study over a reader population), not an open question. Applying G-theory to model raters is also done (arXiv:2507.19980, seven AI raters). The empirical basis was independently unsound: see retraction 16 in the Findings table above. Recorded as prior-art §11 instead, and the open-problem count does not move.

The second one is the instructive one: it would have imported a criticism of a statistic from the paper that refutes the criticism, because the criticism circulates more widely than the measurement does.


What this list is for

Two things, and neither is penance.

A refuted claim is a measurement. Knowing that η inverts on antinoophors, or that a cost-parity gate cannot be met by instructing a budget, is knowledge the programme did not have before, and it was purchased at a price. A file containing only survivors would tell a flattering lie about how the field got here, and would make the same mistakes available to the next person.

And it is an audit surface. The count on the front page is checkable against this table, the table is checkable against the struck-through text, and the struck-through text is checkable against git. A programme whose subject is the gap between confidence and evidence should be the easiest one in the world to catch overstating itself.


This document is licensed CC BY 4.0.

All journal entries · Reference · Noophorics · Repository