audit · prior art

Prior art

Created 2026-07-29, after an external review found that several of this programme's founding novelty claims were false.

This repository requires every number to be reported with its frame. The same standard requires every construct to be reported with its prior art, and PRINCIPIA failed it: it reached for Shannon and Grice and cited not one measurement journal. This file is the correction and the standing obligation.

Every citation below was verified by an agent that located the source independently. Confidence is recorded per item, and the constraints — what each source does not license — are recorded alongside, because that is where the damage happens.


1. Phantom agreement Φ is not ours

The claim "the calibration term — Φ — is the part we have not found elsewhere" is refuted. The gap between believed and actual transfer has been measured, with numbers, in at least four literatures.

sourcewhat it measuredconfidence
Keysar & Henly (2002), Psychological Science 13(3), 207–21240 speaker–listener pairs. Speakers' per-trial judgment matched their own intention on 72% of trials; listeners' actual choice matched on 61%; t(39) = 4.36, p < .001.primary source read
Newton (1990), The Rocky Road from Actions to Intentions, Stanford doctoral dissertation, Ch. 2Tappers estimated listeners would identify ~50% of tapped tunes (range 10–95%); listeners identified 3 of 120, 2.5%.primary source read
Chang, Arora, Lev-Ari, D'Arcy & Keysar (2010), Pediatrics 125(3), 491–496Clinical handoff. The item the sender considered most important was not successfully communicated ~60% of the time, while quality ratings stayed high.abstract read
Endsley (2020), The Divergence of Objective and Subjective Situation Awareness: A Meta-Analysis, JCEDM 14(1), 34–5337 studies carrying both probe-based and self-report measures of situation awareness; the two diverge.abstract read

Their instruments are in places better than ours. Keysar & Henly's headline result is not the 72-vs-61 gap at all — it is a conditional asymmetry: where the addressee had not understood, speakers judged that they had on 46% of trials, against 12% in the other direction, a 34-point gap (t(39) = 6.74, p < .001) present in 80% of individual speakers. Computing that requires pairing each prediction with that trial's outcome. Our single global elicitation of Ĉ cannot compute it at all. That is why falsification criterion 2 was invalid as written.

What these sources do not license

What survives

A claim about coverage, not discovery: we have not found one instrument reporting fidelity, cost, and both parties' calibration against a single stated probe measure — and no measurement of this gap exists where sender and receiver are both language models.

That second half is the programme's distinct object.

Updated 2026-07-31 — it is no longer unmeasured. E-002b completed and measured Φ = +0.2961, 95% CI [+0.2200, +0.3649], between a gpt-oss:120b sender and a gpt-oss:120b receiver over 33 probes, 16 briefs and 1056 probe-party elicitations. As far as this repository has been able to establish, that is the first such measurement between two language models.

Three things must travel with that sentence or it becomes the kind of claim this page exists to prevent:

The claim that survives, in full: the level has been measured once, in the easiest available configuration, and what it mostly shows is that the parties say yes to nearly everything.


2. F* is a normalized-recovery statistic

sourcerelationconfidence
Burns et al., Weak-to-Strong Generalization, arXiv:2312.09390; ICML 2024, PMLR 235:4971–5012Defines performance gap recovered, PGR = (weak-to-strong − weak) / (strong ceiling − weak) — the same shape as F*, sign reversed because D is minimized where performance is maximized.primary source read

The form is generic and predates that paper. Three earlier instances were verified on 2026-07-30 and are recorded in §7. Cohen's κ remains unverified and stays uncited.

Two differences must accompany any statement of the parallel: PGR's terms are task performance against ground truth while F*'s are divergence from the sender; and Burns et al. neither define nor endorse an agreement-normalized PGR — that substitution is ours.


3. Fidelity-versus-correctness was separated in ML first — it is not ours, and it is not ML's

sourcerelationconfidence
Stanton, Izmailov, Kirichenko, Alemi & Wilson (2021), Does Knowledge Distillation Really Work?, NeurIPS 34, 6906–6919Distinguishes fidelity (student–teacher agreement; top-1 agreement, predictive KL) from generalization, and shows good student accuracy does not imply good fidelity.primary source read
Burns et al. (above)High student–supervisor agreement is their failure signal — the imitation mode an auxiliary confidence loss exists to suppress.primary source read

Withdrawn 2026-07-30: the primacy word. Cronbach (1955) separated an accuracy score from an assumed-similarity score, each with its own decomposition, and Edwards et al. (2006) measured similarity and accuracy of team mental models as two quantities and compared them as predictors. Both predate the machine-learning work credited above, which is correctly described and stays. The section's content survives; "first" does not.

E-001's central finding is an independent rediscovery of this split, in a different domain, at the cost of a live run and an amendment.

Not licensed: that Stanton et al. say anything about understanding, meaning or communication (they do not — their evidence is supervised image classification and their prescription is to pursue fidelity harder), or that their result is general (it is regime-dependent: fidelity and generalization are in tension under self-distillation and positively correlated when distilling large ensembles). "Construct failure" is our vocabulary about our metric.

Consequently, PRINCIPIA's claim that knowledge distillation "measures success as task accuracy" was false as written and is corrected in place.


4. L5's human half is established; L5 is not

sourcewhat it measuredconfidence
Carpenter, Wilford, Kornell & Mullaney (2013), Psychon. Bull. Rev. 20(6), 1350–135665-second script held constant, delivery varied. Judgments of learning t(40) = 3.34, d = 1.03; self-rated learning d = 1.78; all four instructor-evaluation items rose. Free recall did not differ. Restudy time equal (1.39 vs 1.43 min, p = .88).primary source read
Deslauriers, McCarty, Miller, Callaghan & Kestin (2019), PNAS 116(39), 19251–19257Randomized crossover: tested learning +0.46 SD, felt learning −0.56 SD, both P < 0.001.primary source read

Neither paper varies narrative against contrastive encoding, so neither addresses L5's refutation condition. Neither measures a sender's belief, so half the Φ construct is untouched. Carpenter's learning result is an accepted null, not a measured decrease, and Deslauriers manipulated active-versus-passive engagement, not fluency.

L5's status stays conjectured. Under this repository's own rules a prior is not a test. What is wrong is the framing "our sharpest conjecture", which is retracted.


5. L6 is not the field's first engineering prescription

sourcewhat it measuredconfidence
Starmer et al. (2014), NEJM 371(19), 1803–1812I-PASS structured handoff across nine pediatric residency programmes, 10 740 admissions: 23% reduction in medical errors (24.5 → 18.8 per 100 admissions) and 30% reduction in preventable adverse events (4.7 → 3.3 per 100), both P < 0.001, without lengthening handoffs.abstract read

Not licensed: that 23% refers to preventable adverse events — that swap appears in AHRQ PSNet's own summary and must not be inherited; that receiver read-back caused the effect (it was an undecomposed five-part bundle with no component-level analysis); that it is causal in the trial sense (pre–post, no concurrent control; six of nine sites improved, one worsened); or that its effect sizes benchmark a prose-versus-constraint experiment, which it never ran. Do not cite SBAR alongside it — SBAR was not verified.

L6 remains a testable claim. It is no longer a primacy claim, and any future experiment should compare against a structured baseline rather than against plain prose, which is a strawman.


6. The standing obligation

A construct enters theory/ with its prior art or it does not enter. The same rule the repository already applies to numbers.

Two operational consequences:

  1. Verify before citing. Every citation on this page was located independently, and the verification caught three misdescriptions in the material it was given — a wrong title attached to a real paper, a pooled statistic presented as a subgroup one, and an outcome swapped for a different outcome in the same study. A fabricated or misdescribed citation in a document about honest measurement is the worst defect available to us.
  2. Record confidence. "Abstract read" and "primary source read" are different epistemic states and are labelled as such. "It circulates widely" is not verification.

The line that must stay sharp: Burns et al. is the one verified source operating on language models, and it measures student–supervisor agreement against ground truth — not either party's belief about the transfer. No verified source has measured a sender's or receiver's claimed agreement in a language model. The fidelity-versus-correctness split is prior art in ML; phantom agreement between language models is not. That, and only that, is what this repository may still call its own — and it may not call it a result until it has one.


7. F* is the judge–advisor statistic under a different distance

sourcewhat it isconfidence
Gino, F. & Moore, D. A. (2007), J. Behavioral Decision Making 20(1), 21–35Publishes WOA = (final estimate − initial estimate) / (advice − initial estimate) verbatim, and traces it to Hell et al. (1988) and Harvey & Fischer (1997)primary source read
Bailey, Leon, Ebner, Moustafa & Weidemann (2022), Current Psychology 42(28), 24516–24541Meta-analysis: pooled WOA = 0.39, 95% CI [0.37, 0.42], k = 346 effect sizes, 129 datasets, N = 17 296primary source read
Yaniv, I. & Kleinberger, E. (2000), OBHDP 83(2), 260–281Publishes the complement, `WOE =a − f/a − i`. Study 1 mean WOE 0.71, so WOA ≈ 0.29primary source read

The correspondence, stated as algebra. Writing D_prior = D(R, B | P) and D_post = D(R, B|m | P), WOA is (D_prior − D_post) / D_prior with the advisor as the reference, the judge's initial estimate as B, the judge's final estimate as B|m, and absolute distance on the line in place of Jensen–Shannon divergence. That is F* exactly as defined in §3, and the estimator of §4.1 only where D_floor = 0.

It is not a generalisation, and the word must not be used. WOA's own instances — scalar magnitude estimates on a continuous scale — lie outside F*'s domain, because §1.2 requires probes with finite discrete answer spaces. Neither contains the other. They are the same functional under different distances. The correspondence is our algebra: no source in this literature frames WOA as a special case of anything, none works with distributions, and none declares a probe measure.

What F* adds: distributions on both sides rather than a point per party; a declared frame (A2), where in JAS the trial is the frame; an explicit finite-sample floor and admissibility gate, which exist because agents are resampleable and human judges are not; and an unclipped negative range — JAS truncates away-from-advice movement to zero, and Bailey et al. state that this truncation "biases the results toward finding evidence for advice-taking."

What F* loses, and this is the uncomfortable direction. JSD is symmetric and non-negative, so a receiver that overshoots the reference is scored identically to one that fell short of it by the same distance. WOA distinguishes them: it exceeds 1 on overshoot. On this axis the older statistic is strictly more informative than ours.

Not licensed

8. Murphy: the skill-score form, and the decomposition we reinvented

sourcewhat it isconfidence
Murphy, A. H. (1988), Monthly Weather Review 116, Eq. (2)SS = (A_f − A_r)/(A_p − A_r) with the reference A_r as a named argument of the score. Murphy calls the form traditional and cites Murphy & Daan (1985)primary source read
Murphy, A. H. (1973), J. Applied Meteorology 12(4), 595–600The reliability / resolution / uncertainty partition of the Brier scoreprimary source read

Murphy (1988) is the precedent for the change this repository is making: in forecast verification the reference is declared as an argument, and which reference you pick changes the score. The anomaly is our current form, which fixes the reference to the sender and never declares it.

Murphy (1973) is the decomposition this repository arrived at independently when it corrected falsification criterion 2 — bias and resolution are different quantities and a mean difference is only the first.

Not licensed: calling it "Murphy's skill score" as an eponym — he presents Eq. (2) as already traditional; and attributing the reliability/resolution partition to the 1988 paper, where it does not appear.

9. Hake's normalized gain, and a criticism that runs backwards

sourcewhat it isconfidence
Hake, R. R. (1998), Am. J. Physics 66(1), 64–74⟨g⟩ = (post − pre)/(100 − pre), over class averages, not per studentprimary source read

The same normalize-the-headroom move as F*, over 62 courses and ~6 000 students.

Not licensed, and this one nearly went in backwards: the sentence "normalized gain is known to correlate with pretest score" was drafted for this page and is contradicted by the paper cited for it — Hake reports r = +0.02 across 62 courses, and that near-zero correlation is his central justification for the measure. It also may not be presented as a per-student score.

10. The vanishing denominator is not our problem alone, and nobody has solved it

F* is undefined when D_prior → D_floor: there is no gap to close. The judge–advisor literature has the identical defect and has had it in print since at least 2006 — Gino & Moore state that WOA "yields undefined values when the advice is equal to the judge's initial estimate."

Their resolution is exclusion plus truncation: drop the undefined trials, clip the rest into [0, 1]. Ours is an ε-gate (exclusion) and min(1.0, …) (truncation). We arrived at the same two workarounds.

Bailey et al.'s meta-analysis states that the truncation biases estimates upward. Ours is the same truncation.

Recorded as Problem 12. Citing JAS as precedent for a fix is forbidden: they have the defect, not a solution.


11. The reader is a facet, and this programme fixed it without saying so

A2 indexes every quantity by a probe measure P and a declared reference R. It does not index by the agent that reads P, and the reporting standard in definitions §7 files sender and receiver identity as a disclosure field — something you state about a measurement, not something the measurement is indexed by. Measurement theory has had a name for that choice, and a price for it, since the 1970s.

sourcewhat it establishesconfidence
Brennan, R. L. (2003), Coefficients and Indices in Generalizability Theory, CASMA Research Report No. 1, University of Iowa, p. 27Every condition of measurement is a facet — persons, items, raters, occasions — and a facet may be random or fixed. Verbatim: "in any generalizability analysis there must be at least one random facet for the analysis to be meaningful. If all facets were fixed, then no generalization is involved, and all error variances are zero, be definition." (the typo is the original's)primary source read — PDF obtained and the passage extracted verbatim
Dorans, N. J. & Holland, P. W. (1992), DIF Detection and Description, ETS RR-92-10 / ERIC ED387526, pp. 5–8DIF is not a rate difference. Verbatim: "DIF is an unexpected difference among groups of examinees who are supposed to be comparable with respect to the attribute measured by the item and test on which it appears", and "Simpson's paradox (Simpson, 1951) illustrates why we should compare the comparable, as is done in DIF analyses." An unconditioned difference between groups of unequal ability is impact, a different quantity.primary source read — PDF obtained and the passages extracted verbatim
Song, Lee & Jiao (2025), Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory, arXiv:2507.19980A generalizability study with two human and seven AI raters as an explicit rater facet. Applying G-theory to model readers is done work.abstract read

What this means for this repository, stated plainly. Every characterization MERIDIAN-IX32 carries — which probes discriminate, the per-probe rate, the headroom against E-002c's gate — was estimated with the reader held fixed at gpt-oss:120b. In G-theory's vocabulary that is a fixed facet, and Brennan's sentence is not a caution but a theorem: with the facet fixed there is no generalization over readers and the corresponding error variance is zero by construction, not by measurement. The remedy is equally standard — a D-study over a population of readers — and it is a design cost, not a research problem.

Not licensed. That the field has studied this programme's quantities: it has not. G-theory decomposes the variance of a score into facets; it says nothing about F*, about Φ, or about a sender's belief regarding what transferred. What is prior art is the structural point — that an instrument read through one rater cannot separate its own properties from that rater's — and it is prior art completely. What is not prior art is any noophoric quantity computed under that structure.

Also not licensed: reading the two-reader result as differential item functioning. gpt-oss and qwen3.5:35b differ in overall ability on this measure — a fitted gap of 2.20 logits — and the comparison was never conditioned on it. By Dorans & Holland that is impact, and Simpson's paradox is the named hazard. Recorded as retraction 16.


This document is licensed CC BY 4.0.

All journal entries · Reference · Noophorics · Repository