Reference
Every quantity, law, experiment, withdrawn claim and open problem in the programme. This page is generated from the files that define each of them — nothing here is typed twice, so it cannot drift from the theory the way a hand-kept index would.
Symbols19
From the symbol table in lexicon.md.
| Symbol | Reads as | Defined in |
|---|---|---|
| π | a probe | §1.2 |
| P | probe measure | §1.3 |
| m | the message / artifact | §1.4 |
| B\|m | receiver after conditioning on m | §1.4 |
| d(A,B\|π) | per-probe divergence | §2.1 |
| D(A,B\|P) | divergence over P | §2.2 |
| Â | observed agreement rate | §2.3 |
| D_self | self-divergence | §3.1 |
| D_floor | noise floor | §3.2 |
| F*_R | transfer fidelity | §4.1 |
| C(m) | message cost | §4.2 |
| η | noophoric efficiency | §4.3 |
| Ĉ | claimed agreement rate | §5 |
| Φ | phantom agreement | §5 |
| β | calibration slope | §5.1 |
| K_R | channel capacity | §6.1 |
| U | residual | §6.2 |
| R | reference disposition | §4.1.1 |
| V_λ | net value | §4.3 |
Terms29
Struck text is a withdrawn claim kept in place, never deleted.
- noophorics
- the quantitative science of transferring understanding across systems with non-identical priors. From νόος (mind) + φορά (carrying).
- noophor
- a single act of transfer: sender, artifact, receiver.
- antinoophor
- a noophor with
F* < 0. A message that leaves the receiver further from the sender than before it was sent. - noophoric act
- the triple
(A, m, B). Synonym for noophor; used when the components matter. - criterion-bearing / criterion-free
- the two regimes a probe measure can be in (v0.4). *Criterion-bearing*: an answer exists independently of the sender, so a key or an adjudicator panel is an admissible reference and the word *understanding* is licensed. *Criterion-free*: the sender is the criterion by construction — a preference, a house style, a judgment call whose owner defines the right answer — so
R = senderis correct rather than tolerated, and there is no sender error available to be replicated. Key *absence* does not move a measure into the second regime: that is an epistemic fact about the experimenter, not an ontological one about the probe. - agent
- any system that maps a probe to a distribution over answers. Defined entirely by its answer distributions; internals are irrelevant.
- sender / receiver
- the two roles in a noophor. Roles, not properties: the same system is routinely both.
- probe
- a decidable decision whose answer depends on the understanding being transferred. *Decidable* means the answer space is finite and discrete.
- probe measure (
P) - a distribution over probes. The noophoric frame of reference. No quantity is defined without one.
- admissible probe measure
- one on which the parties actually disagree before transfer. Measuring fidelity where agreement already exists measures nothing.
- divergence (
D) - expected Jensen–Shannon divergence between the agents' answer distributions over
P, in bits. Bounded[0, 1]. - agreement rate (
Â) - fraction of probes on which the agents' modal answers match. Coarser than
D, but the quantity parties can actually estimate about themselves. - self-divergence (
D_self) - an agent's divergence from itself under independent resampling. Agents are stochastic; this is the measure of it.
- noise floor (
D_floor) mean self-divergence of the two agents; the irreducible divergence between perfectly aligned parties. Retracted in v0.3 (retraction 2): two perfectly aligned stochastic agents have identical true distributions, so their true divergence is exactly zero and there is nothing irreducible.D_flooris finite-sample estimator bias at a statedn, obtained by a permutation null over the pooled draws. It is a property of the measurement, not of the agents, and it belongs in the estimator rather than in the definition. Not correcting for it is still the most common error.- transfer fidelity (
F*_R) - the fraction of the closable gap toward a declared reference
Rthat the message closed, floor-corrected.1= fully closed,0= no effect,< 0= antinoophor. Never reportable without itsR(v0.4).F*_{R=A}is identically the pre-v0.4 quantity. - cost (
C) - the price of the artifact, in the receiver's tokens unless stated otherwise. The receiver pays to read it.
- noophoric efficiency (
η) F*/C, understanding per unit cost, and valid only whereF* ≥ 0.The quantity engineering should optimize.A ratio with a signed numerator is not an ordering (retraction 3): atF* = −1a 100-token antinoophor scores −10.00 and an 800-token one −1.25, ranking the costlier failure higher. UseV_λwhen the sign is unknown.- net value (
V_λ) F*_R − λ·C. Monotone in both arguments at every sign, so it orders messages the way the field means to.λis the declared exchange rate between fidelity and a token; sweeping it traces the frontier, which makesV_λandK_R(C)the same object seen twice.- channel capacity between minds (
K_R) sup_m F*_R(m). The best fidelity any message could achieve at unbounded cost, toward a declared reference. Estimated as a lower boundK̂over a stated search budget, by sample-splitting — a max-over-search estimate is a winner's curse and overstates.- residual (
U) 1 − K_R. The untransferable remainder toward a declared reference. Axiom A3 asserts it is nonzero.Written— renamed because v0.4 gaveRbefore 2026-07-31Rto the reference disposition without checking the symbol was free, and for a day this glossary listedRas the residual directly above a section definingRas the reference.- phantom agreement (
Φ) - claimed agreement rate minus observed agreement rate.
Φ > 0is a shared illusion of successful transfer: both parties believe it landed, probes say otherwise. The field's central pathology. - calibration slope (
β) d(claimed)/d(observed): how far a party's claim moves when the outcome moves.1is calibrated in the responsive sense,0is inert, below0is anti-calibrated. Independent of [[phantom agreement]]: a party claiming the long-run mean every time scoresΦ = 0andβ = 0, and is maximally uninformative. ReportingΦwithoutβis the same error as reporting bias without resolution. Report it per party and report the attenuation-corrected value beside the raw one, never instead of it. Pre-registered and measured in E-002c: sender−0.02, receiver+0.28.- invariant core
- the part of an understanding that survives arbitrarily many chained transfers without loss. Conjectured (L4) to be constraints and prohibitions rather than descriptions.
- contrastive encoding
- a message specifying where the parties would diverge: boundaries, exclusions, edge cases, what *not* to do.
- declarative encoding
- a message describing what the sender understands. The default form of nearly every handoff written today.
- pre-registration
- committing an experiment's hypothesis and analysis plan before any data exists. The git history is the record.
- probe test
- running a probe measure after a transfer to measure what landed. The noophoric analogue of replication.
- ablation ladder
- truncating a message to successive cost levels and measuring the resulting
F*(C)curve. - reconstructive test
- asking whether the receiver can generate a message that achieves comparable fidelity with a *third* party. Tests whether understanding transferred deeply enough to be re-transmitted.
Conjectural laws6
A law enters the record only with a refutation condition attached, and leaves it never. From theory/laws.md.
| Law | Status | Refuted if | |
|---|---|---|---|
| L1 | Law of diminishing noophoric return | conjectured | F*(C) is linear over a substantial range, or reaches 1.0 (floor-corrected) for arbitrary sender/receiver pairs. |
| L2 | Law of asymmetry | conjectured · *headline form withdrawn as tautological* | across pairs where domain prior and general capability point in opposite directions, the sign of the asymmetry follows capability, or follows neither at better than chance. |
| L3 | Law of prior overlap | conjectured · *operationalization incomplete* | a non-circular overlap measure exists and fails to predict K. |
| L4 | The curse of the summary | conjectured · *restated after the original was found ill-typed* | chain fidelity is not monotone (L4a); log F* is not approximately linear over the positive range (L4b); or constraint-form and description-form content show decay slopes that are statistically indistinguishable (L4c). |
| L5 | Fluency inflates phantom agreement | conjectured | at equal cost, narrative and contrastive encodings produce statistically indistinguishable Φ, or narrative produces *lower* Φ. |
| L6 | Optimal encoding is contrastive, not declarative | conjectured | contrastive encodings show η ≤ declarative encodings at equal cost, across domains. |
Experiments8
Status is read from which files exist in each directory — a
VOID.md, a FINDINGS.md, a
PREREGISTRATION.md — not from a label anybody maintains.
| Id | Status | Results | Notes |
|---|---|---|---|
| E-001-fluency-cost | findings | 1 | — |
| E-001b-fluency-factorial | void | — | defect recorded |
| E-001c-fluency-length-controlled | void | 2 | defect recorded |
| E-002-phantom-agreement | void | 1 | — |
| E-002b-phantom-agreement-ladder | findings | 1 | — |
| E-002c-calibration-slope | findings | 3 | defect recorded |
| E-004-disagreement-detector | void | 3 | interrupted, resumed, an arm was blocked |
| E-006-ablation-ladder | void | — | — |
Withdrawn claims21
Ours, with what killed each one. Full index at retractions.
| # | Claim | Killed by |
|---|---|---|
| 1 | A 113-token lookup table reaching F* = 1. Restated for held-out probes under bounded cost. | |
| 2 | D_floor is "the irreducible divergence caused by the parties' own stochasticity", and belongs inside the *definition* of F* | Two perfectly aligned stochastic agents have identical true distributions, so their true JSD is exactly zero. The floor is estimator bias, and a fidelity that changes when you sample more is not well-posed. |
| 3 | η = F*/C is "the quantity engineering should optimize" | A ratio with a signed numerator is not an ordering. At F* = −1, a 100-token antinoophor scores −10.00 and an 800-token one −1.25, so the message that spends eight times as much to do the same damage ranks higher. Replaced by V_λ = F* − λC. |
| 4 | Φ ≈ 0 means the pathology does not exist" | Bias and resolution are independent. A party predicting 0.70 on every probe and averaging 0.70 has Φ = 0 and no ability to say which probes it got wrong; it is maximally pathological and the criterion scored it as our refutation. |
| 5 | Tautological as stated. | |
| 6 | Ill-typed: it multiplied fractions of different prior gaps, and two antinoophors composed to a positive product. Measured: hops of −0.629 and −1.000 multiply to +0.629. Restated as L4a/L4b/L4c. | |
| 7 | I-PASS, deployed and outcome-measured since 2014 across nine programmes and 10 740 admissions. | |
| 8 | The human half is Carpenter et al. (2013) and Deslauriers et al. (2019). Status stays conjectured — a prior is not a test — but the framing was ours to lose. | |
| 9 | Φ is "the part we have not found elsewhere", and is what "everyone had felt, nobody had weighed" | Keysar & Henly (2002), Newton (1990), Chang et al. (2010), Endsley (2020). It was weighed in 1990, and their instruments are in places better than ours. |
| 10 | Stanton et al. (2021) define fidelity separately from generalization and show accuracy does not imply it — E-001's construct failure, from a NeurIPS abstract, five years early. | |
| 11 | Cronbach (1955) separated an accuracy score from an assumed-similarity score, each with its own decomposition; Edwards et al. (2006) measured team mental-model similarity and accuracy as two quantities and compared them as predictors. Both predate the ML work. The content survives; the primacy word does not. | |
| 12 | E-002c committed the quantity before collecting and measured β = +0.1299, CI [+0.047, +0.223], which clears zero. Confidence does respond, at about an eighth of the rate calibration requires. The direction replicated and the magnitude did not. The unresponsive party is the sender alone, β = −0.02 with its interval spanning zero — a narrower claim than the one withdrawn, and a sharper one. | |
| 13 | interaction probes diverge at 230 words" — stated in Problem 15 from six messages, and read as the property that only interaction probes survive saturation | Widening the same measurement to twelve messages across all four cells found three: M02, M12, M29. A small-sample zero, killed within a day by the instrument that produced it. What survives is the proportion — 21 of 24 divergence events on 9 of 34 probes — and the divergence rate the specification is costed against, 0.194 against 0.204, which the widening left standing. |
| 14 | MERIDIAN-IX32 reaches 3.33 diverged probes per message and clears E-002c's outcome-variation gate" — the candidate probe measure built to repair Problem 15 | The admission gate behind that number had been applied once. Applied five times, admitting a probe only if it returns the key at margin ≥ 8 in every pass, it rejects four probes, and the measure on the 28 survivors gives 2.00 per message — below the gate. Worse for the repair than for the claim: all four rejected probes are among the nine that ever discriminated, and none of the twenty-three that never discriminated was rejected, p = 0.0035. The headroom was borrowed from probes that are not stable observables. |
| 15 | p computed over MERIDIAN-IX32's probes — p = 0.0035 and p = 0.0138 for the anti-correlation between discriminating power and modal stability, and p = 0.0339 for the design rule that earned X17–X32 | The measure is not 32 independent probes. X17 and X18 differ by one token — "Combined figure count is 70" against "71" — and carry opposite keys; so do X21/X22 (8 → 7) and X24/X25 (30 → 21). Single-link clustering of the prompts gives 9 clusters at similarity 0.80–0.85, the largest holding 11 probes, and all instability falls inside 2 of them. Every one of these tests treats near-duplicate rows as independent draws. There is no corrected number to substitute, which is the point: the cluster-level p ranges from 0.111 at threshold 0.80 to 0.006 at 0.95, so the result is a function of a clustering knob rather than of the data. What survives is at the draw level and is stronger — of 1 440 gpt-oss:120b sender draws, 19 are non-key and all 19 fall on R5-tagged probes, a label committed 2026-08-10 09:29:56, nine hours before the first IX32 sender pass and therefore prior to every instability datum. |
| 16 | MERIDIAN-IX32 measures model-independent transfer loss on 2 of 32 probes *because* the other seven are one reader's uncertainty | Three independent grounds. (a) The statistic was not like-for-like. "Six of the seven sit at gpt-oss's wobble points" compares gpt-oss's *minimum over four sender passes* against qwen's *single* pass; more passes means more chances to show a low margin. Counted one pass each it is 2 of 7. (b) No reader-specific term is needed. A model with one difficulty per probe and a single reader-ability gap — no reader×probe interaction whatever — fits at deviance 10.54 on 8 df, p = 0.229, with the gap at 2.20 logits. The strict-subset structure is what that model predicts anyway. (c) It is impact, not differential functioning. Dorans & Holland (1992): DIF requires comparing examinees "supposed to be comparable with respect to the attribute measured"; an unconditioned difference between groups of unequal ability is *impact*, and Simpson's paradox is the named hazard. The comparison was never conditioned on the 2.20-logit gap. The counts survive — qwen loses 2 of 32, gpt-oss 9, qwen's set a strict subset — but *why* is open, and was published as answered. |
| 17 | +0.228 / +0.456 / +0.606 / +0.848, and the three briefs reported at exactly 1.0000 | The noise floor was computed on a different pair from the comparison it corrects. runner.py:222 takes the permutation floor between the sender and PRIOR; line 241 then divides every sender-versus-receiver fidelity by it. D_floor is estimator bias *for the pair being compared*, and these parties are not alike — the sender is a point mass on 33 of 33 probes, PRIOR on 4 — so the mismatched floor is 0.0417 against matched floors of 0.0076–0.0226. Every fidelity was inflated, and because the matched floor falls as cost rises the inflation steepened the ladder as well as lifting it: corrected, the column is +0.217 / +0.424 / +0.560 / +0.784, the climb is +0.57 not +0.62, and the three briefs at the min(1.0, ·) cap become 0.912 / 0.915 / 0.953 — no brief on the ladder reaches 1.0. β, the registered primary quantity, is a slope of claim on *observed agreement* and never reads this column; §3's finding is untouched. |
| 18 | A second model satisfies it. The same script, specification, calibrated instruction and band, changing only the model: qwen3.5:35b puts cell A in band 11 of 12 against gpt-oss:120b's 0 of 12, and cell B 12 of 12 against 5. gpt-oss and 197 for qwen — above the 231 ceiling for one, 34 words below it for the other.gpt-oss and 23 of 24 on qwen, against terse's 23 of 24 on both — so on qwen there is no fluency-axis length effect at all. The void itself stands: E-001c ran on gpt-oss, that model could not compose inside its own band, and the experiment correctly died. What is withdrawn is the generalisation from one generator to the register. L5/L6 are not revived — this measures the band filter only, and E-001c's gate was the band *and* two blind raters, whose intersection is where the 0 came from. | |
| 19 | gpt-oss's fluent floor sits above [the band ceiling]" — the gloss under retraction 18's own evidence table, published in the results file and repeated in row 18 above | The table it sits under refutes it. RESULTS-qwen-floor.md reports gpt-oss's fluent floor, minimum over cells A and B, as 223. The ceiling is 231. 223 is eight words below it. The claim is true only of cell A (floor 232, 0/12 in band); cell B's floor is 223 and it lands in band 5 of 12 — so gpt-oss reaches the band under a *fluency instruction*, in the contrastive cell, and what it cannot do is fluent × declarative specifically, which is the cell E-001c's VOID names in its gate. That is a statement about length and not about register: this run measures the band alone, and E-001c's VOID records that of the two live cell-A messages that did reach the band, *neither was judged fluent prose*. qwen's "34 words below" is correct. Retraction 18 is untouched: it rests on the in-band counts — 11 of 12 against 0 of 12 in cell A, 12 of 12 against 5 of 12 in cell B — not on this sentence. A wrong gloss was carried on top of a right finding for six days, into the ledger and onto the front page, because the number under it was never subtracted from the number beside it. |
| 20 | claude-* model has ever read RIVERSIDE-30" and "It has never been run on RIVERSIDE-30" — the founding premise of the 2026-09-01 crossover prediction, and with it the claim that the measurement was blocked on an absent API key | It had been run on 2026-08-03, and the file that proves it is the file the same paragraph cites. E-004 measured claude-opus-4-8 against RIVERSIDE-30@2e6afe2f3c92 — the identical measure hash — at 0.733, beside gpt-oss's 0.967, in E-004-20260803T161541Z.json, committed 801ac86. E-001c's PARAMETERS had quoted that very figure since 2026-08-03. The prediction is scored and both its sharp clauses hold — below 1.000, and acc(claude) − acc(gpt-oss) negative on both domains (−0.233, −0.091) — so *no crossover* now rests on four models rather than three, and rested on enough a month before it was argued. What is withdrawn is the reason the question was thought open: the 2026-09-01 → 09-06 arc, 42 GB of weights and 1 280 probe calls, was launched to settle something already settled and filed. The exposure claim in the same file survives — a stateless subject is not a contaminated agent. |
| 21 | theory/laws.md, on the front page, and in three probes/ documents | E-004 registered no such gate. Its pre-registration §5.1 sets each model's accuracy > 0.60 on each measure, and its own void note records that *"every model cleared the 0.60 accuracy floor"*. On E-004's real gate llama3.3:70b passes MERIDIAN-34 at 0.824 and fails only RIVERSIDE-30 — so a reader who audited the citation found the opposite of what the sentence claimed, on one of the two domains. The 0.90 threshold is real and is the sender accuracy vs key > 0.90 gate registered in E-001b, E-001c, E-002b and E-002c; llama3.3 fails *that* on both. The conclusion survives, its authority did not: the same shape as retraction 16 in the Findings table above, where the counts held and the reason given for them did not. |
Open problems15
| # | Problem |
|---|---|
| 1 | The residual characterization problem |
| 2 | Non-circular prior overlap |
| 3 | The capacity theorem |
| 4 | Optimal encoding search |
| 5 | The asymmetry law |
| 6 | Probe measure generalization |
| 7 | Phantom agreement mechanics |
| 8 | The invariant core |
| 9 | Self-transfer |
| 10 | The multi-party problem |
| 11 | Post-transfer admissibility |
| 12 | The vanishing denominator |
| 13 | A hypothesis's reported statistic and its tested statistic are not checked against each other |
| 14 | When is a modal answer over n draws a stable observable? |
| 15 | Φ has no belief component where the manipulation has to live |