Reference

Every quantity, law, experiment, withdrawn claim and open problem in the programme. This page is generated from the files that define each of them — nothing here is typed twice, so it cannot drift from the theory the way a hand-kept index would.

Symbols19

From the symbol table in lexicon.md.

SymbolReads asDefined in
πa probe§1.2
Pprobe measure§1.3
mthe message / artifact§1.4
B\|mreceiver after conditioning on m§1.4
d(A,B\|π)per-probe divergence§2.1
D(A,B\|P)divergence over P§2.2
Âobserved agreement rate§2.3
D_selfself-divergence§3.1
D_floornoise floor§3.2
F*_Rtransfer fidelity§4.1
C(m)message cost§4.2
ηnoophoric efficiency§4.3
Ĉclaimed agreement rate§5
Φphantom agreement§5
βcalibration slope§5.1
K_Rchannel capacity§6.1
Uresidual§6.2
Rreference disposition§4.1.1
V_λnet value§4.3

Terms29

Struck text is a withdrawn claim kept in place, never deleted.

noophorics
the quantitative science of transferring understanding across systems with non-identical priors. From νόος (mind) + φορά (carrying).
noophor
a single act of transfer: sender, artifact, receiver.
antinoophor
a noophor with F* < 0. A message that leaves the receiver further from the sender than before it was sent.
noophoric act
the triple (A, m, B). Synonym for noophor; used when the components matter.
criterion-bearing / criterion-free
the two regimes a probe measure can be in (v0.4). *Criterion-bearing*: an answer exists independently of the sender, so a key or an adjudicator panel is an admissible reference and the word *understanding* is licensed. *Criterion-free*: the sender is the criterion by construction — a preference, a house style, a judgment call whose owner defines the right answer — so R = sender is correct rather than tolerated, and there is no sender error available to be replicated. Key *absence* does not move a measure into the second regime: that is an epistemic fact about the experimenter, not an ontological one about the probe.
agent
any system that maps a probe to a distribution over answers. Defined entirely by its answer distributions; internals are irrelevant.
sender / receiver
the two roles in a noophor. Roles, not properties: the same system is routinely both.
probe
a decidable decision whose answer depends on the understanding being transferred. *Decidable* means the answer space is finite and discrete.
probe measure (P)
a distribution over probes. The noophoric frame of reference. No quantity is defined without one.
admissible probe measure
one on which the parties actually disagree before transfer. Measuring fidelity where agreement already exists measures nothing.
divergence (D)
expected Jensen–Shannon divergence between the agents' answer distributions over P, in bits. Bounded [0, 1].
agreement rate (Â)
fraction of probes on which the agents' modal answers match. Coarser than D, but the quantity parties can actually estimate about themselves.
self-divergence (D_self)
an agent's divergence from itself under independent resampling. Agents are stochastic; this is the measure of it.
noise floor (D_floor)
mean self-divergence of the two agents; the irreducible divergence between perfectly aligned parties. Retracted in v0.3 (retraction 2): two perfectly aligned stochastic agents have identical true distributions, so their true divergence is exactly zero and there is nothing irreducible. D_floor is finite-sample estimator bias at a stated n, obtained by a permutation null over the pooled draws. It is a property of the measurement, not of the agents, and it belongs in the estimator rather than in the definition. Not correcting for it is still the most common error.
transfer fidelity (F*_R)
the fraction of the closable gap toward a declared reference R that the message closed, floor-corrected. 1 = fully closed, 0 = no effect, < 0 = antinoophor. Never reportable without its R (v0.4). F*_{R=A} is identically the pre-v0.4 quantity.
cost (C)
the price of the artifact, in the receiver's tokens unless stated otherwise. The receiver pays to read it.
noophoric efficiency (η)
F*/C, understanding per unit cost, and valid only where F* ≥ 0. The quantity engineering should optimize. A ratio with a signed numerator is not an ordering (retraction 3): at F* = −1 a 100-token antinoophor scores −10.00 and an 800-token one −1.25, ranking the costlier failure higher. Use V_λ when the sign is unknown.
net value (V_λ)
F*_R − λ·C. Monotone in both arguments at every sign, so it orders messages the way the field means to. λ is the declared exchange rate between fidelity and a token; sweeping it traces the frontier, which makes V_λ and K_R(C) the same object seen twice.
channel capacity between minds (K_R)
sup_m F*_R(m). The best fidelity any message could achieve at unbounded cost, toward a declared reference. Estimated as a lower bound over a stated search budget, by sample-splitting — a max-over-search estimate is a winner's curse and overstates.
residual (U)
1 − K_R. The untransferable remainder toward a declared reference. Axiom A3 asserts it is nonzero. Written R before 2026-07-31 — renamed because v0.4 gave R to the reference disposition without checking the symbol was free, and for a day this glossary listed R as the residual directly above a section defining R as the reference.
phantom agreement (Φ)
claimed agreement rate minus observed agreement rate. Φ > 0 is a shared illusion of successful transfer: both parties believe it landed, probes say otherwise. The field's central pathology.
calibration slope (β)
d(claimed)/d(observed): how far a party's claim moves when the outcome moves. 1 is calibrated in the responsive sense, 0 is inert, below 0 is anti-calibrated. Independent of [[phantom agreement]]: a party claiming the long-run mean every time scores Φ = 0 and β = 0, and is maximally uninformative. Reporting Φ without β is the same error as reporting bias without resolution. Report it per party and report the attenuation-corrected value beside the raw one, never instead of it. Pre-registered and measured in E-002c: sender −0.02, receiver +0.28.
invariant core
the part of an understanding that survives arbitrarily many chained transfers without loss. Conjectured (L4) to be constraints and prohibitions rather than descriptions.
contrastive encoding
a message specifying where the parties would diverge: boundaries, exclusions, edge cases, what *not* to do.
declarative encoding
a message describing what the sender understands. The default form of nearly every handoff written today.
pre-registration
committing an experiment's hypothesis and analysis plan before any data exists. The git history is the record.
probe test
running a probe measure after a transfer to measure what landed. The noophoric analogue of replication.
ablation ladder
truncating a message to successive cost levels and measuring the resulting F*(C) curve.
reconstructive test
asking whether the receiver can generate a message that achieves comparable fidelity with a *third* party. Tests whether understanding transferred deeply enough to be re-transmitted.

Conjectural laws6

A law enters the record only with a refutation condition attached, and leaves it never. From theory/laws.md.

LawStatusRefuted if
L1Law of diminishing noophoric returnconjecturedF*(C) is linear over a substantial range, or reaches 1.0 (floor-corrected) for arbitrary sender/receiver pairs.
L2Law of asymmetryconjectured · *headline form withdrawn as tautological*across pairs where domain prior and general capability point in opposite directions, the sign of the asymmetry follows capability, or follows neither at better than chance.
L3Law of prior overlapconjectured · *operationalization incomplete*a non-circular overlap measure exists and fails to predict K.
L4The curse of the summaryconjectured · *restated after the original was found ill-typed*chain fidelity is not monotone (L4a); log F* is not approximately linear over the positive range (L4b); or constraint-form and description-form content show decay slopes that are statistically indistinguishable (L4c).
L5Fluency inflates phantom agreementconjecturedat equal cost, narrative and contrastive encodings produce statistically indistinguishable Φ, or narrative produces *lower* Φ.
L6Optimal encoding is contrastive, not declarativeconjecturedcontrastive encodings show η ≤ declarative encodings at equal cost, across domains.

Experiments8

Status is read from which files exist in each directory — a VOID.md, a FINDINGS.md, a PREREGISTRATION.md — not from a label anybody maintains.

IdStatusResultsNotes
E-001-fluency-costfindings1
E-001b-fluency-factorialvoiddefect recorded
E-001c-fluency-length-controlledvoid2defect recorded
E-002-phantom-agreementvoid1
E-002b-phantom-agreement-ladderfindings1
E-002c-calibration-slopefindings3defect recorded
E-004-disagreement-detectorvoid3interrupted, resumed, an arm was blocked
E-006-ablation-laddervoid

Withdrawn claims21

Ours, with what killed each one. Full index at retractions.

#ClaimKilled by
1A3 — no bounded message closes an arbitrary prior gapA 113-token lookup table reaching F* = 1. Restated for held-out probes under bounded cost.
2D_floor is "the irreducible divergence caused by the parties' own stochasticity", and belongs inside the *definition* of F*Two perfectly aligned stochastic agents have identical true distributions, so their true JSD is exactly zero. The floor is estimator bias, and a fidelity that changes when you sample more is not well-posed.
3η = F*/C is "the quantity engineering should optimize"A ratio with a signed numerator is not an ordering. At F* = −1, a 100-token antinoophor scores −10.00 and an 800-token one −1.25, so the message that spends eight times as much to do the same damage ranks higher. Replaced by V_λ = F* − λC.
4Falsification criterion 2 — "Φ ≈ 0 means the pathology does not exist"Bias and resolution are independent. A party predicting 0.70 on every probe and averaging 0.70 has Φ = 0 and no ability to say which probes it got wrong; it is maximally pathological and the criterion scored it as our refutation.
5L2 headline formTautological as stated.
6L4 — fidelity is multiplicative along a chainIll-typed: it multiplied fractions of different prior gaps, and two antinoophors composed to a positive product. Measured: hops of −0.629 and −1.000 multiply to +0.629. Restated as L4a/L4b/L4c.
7L6 is "the field's first engineering prescription"I-PASS, deployed and outcome-measured since 2014 across nine programmes and 10 740 admissions.
8L5 is "our sharpest conjecture"The human half is Carpenter et al. (2013) and Deslauriers et al. (2019). Status stays conjectured — a prior is not a test — but the framing was ours to lose.
9Φ is "the part we have not found elsewhere", and is what "everyone had felt, nobody had weighed"Keysar & Henly (2002), Newton (1990), Chang et al. (2010), Endsley (2020). It was weighed in 1990, and their instruments are in places better than ours.
10Knowledge distillation "measures success as task accuracy"Stanton et al. (2021) define fidelity separately from generalization and show accuracy does not imply it — E-001's construct failure, from a NeurIPS abstract, five years early.
11"Fidelity-versus-correctness was separated in ML first"Cronbach (1955) separated an accuracy score from an assumed-similarity score, each with its own decomposition; Edwards et al. (2006) measured team mental-model similarity and accuracy as two quantities and compared them as predictors. Both predate the ML work. The content survives; the primacy word does not.
12"The parties' confidence is very nearly unresponsive to how much actually transferred" — E-002b's headline pull-quote, labelled post-hoc when madeE-002c committed the quantity before collecting and measured β = +0.1299, CI [+0.047, +0.223], which clears zero. Confidence does respond, at about an eighth of the rate calibration requires. The direction replicated and the magnitude did not. The unresponsive party is the sender alone, β = −0.02 with its interval spanning zero — a narrower claim than the one withdrawn, and a sharper one.
13"Zero of the 25 non-interaction probes diverge at 230 words" — stated in Problem 15 from six messages, and read as the property that only interaction probes survive saturationWidening the same measurement to twelve messages across all four cells found three: M02, M12, M29. A small-sample zero, killed within a day by the instrument that produced it. What survives is the proportion — 21 of 24 divergence events on 9 of 34 probes — and the divergence rate the specification is costed against, 0.194 against 0.204, which the widening left standing.
14"MERIDIAN-IX32 reaches 3.33 diverged probes per message and clears E-002c's outcome-variation gate" — the candidate probe measure built to repair Problem 15The admission gate behind that number had been applied once. Applied five times, admitting a probe only if it returns the key at margin ≥ 8 in every pass, it rejects four probes, and the measure on the 28 survivors gives 2.00 per message — below the gate. Worse for the repair than for the claim: all four rejected probes are among the nine that ever discriminated, and none of the twenty-three that never discriminated was rejected, p = 0.0035. The headroom was borrowed from probes that are not stable observables.
15Every Fisher exact p computed over MERIDIAN-IX32's probes — p = 0.0035 and p = 0.0138 for the anti-correlation between discriminating power and modal stability, and p = 0.0339 for the design rule that earned X17X32The measure is not 32 independent probes. X17 and X18 differ by one token — "Combined figure count is 70" against "71" — and carry opposite keys; so do X21/X22 (8 → 7) and X24/X25 (30 → 21). Single-link clustering of the prompts gives 9 clusters at similarity 0.80–0.85, the largest holding 11 probes, and all instability falls inside 2 of them. Every one of these tests treats near-duplicate rows as independent draws. There is no corrected number to substitute, which is the point: the cluster-level p ranges from 0.111 at threshold 0.80 to 0.006 at 0.95, so the result is a function of a clustering knob rather than of the data. What survives is at the draw level and is stronger — of 1 440 gpt-oss:120b sender draws, 19 are non-key and all 19 fall on R5-tagged probes, a label committed 2026-08-10 09:29:56, nine hours before the first IX32 sender pass and therefore prior to every instability datum.
16"Seven of the nine discriminating probes measured the reader, not the transfer" — and its corollary that MERIDIAN-IX32 measures model-independent transfer loss on 2 of 32 probes *because* the other seven are one reader's uncertaintyThree independent grounds. (a) The statistic was not like-for-like. "Six of the seven sit at gpt-oss's wobble points" compares gpt-oss's *minimum over four sender passes* against qwen's *single* pass; more passes means more chances to show a low margin. Counted one pass each it is 2 of 7. (b) No reader-specific term is needed. A model with one difficulty per probe and a single reader-ability gap — no reader×probe interaction whatever — fits at deviance 10.54 on 8 df, p = 0.229, with the gap at 2.20 logits. The strict-subset structure is what that model predicts anyway. (c) It is impact, not differential functioning. Dorans & Holland (1992): DIF requires comparing examinees "supposed to be comparable with respect to the attribute measured"; an unconditioned difference between groups of unequal ability is *impact*, and Simpson's paradox is the named hazard. The comparison was never conditioned on the 2.20-logit gap. The counts survive — qwen loses 2 of 32, gpt-oss 9, qwen's set a strict subset — but *why* is open, and was published as answered.
17E-002c's published per-rung fidelity column — +0.228 / +0.456 / +0.606 / +0.848, and the three briefs reported at exactly 1.0000The noise floor was computed on a different pair from the comparison it corrects. runner.py:222 takes the permutation floor between the sender and PRIOR; line 241 then divides every sender-versus-receiver fidelity by it. D_floor is estimator bias *for the pair being compared*, and these parties are not alike — the sender is a point mass on 33 of 33 probes, PRIOR on 4 — so the mismatched floor is 0.0417 against matched floors of 0.0076–0.0226. Every fidelity was inflated, and because the matched floor falls as cost rises the inflation steepened the ladder as well as lifting it: corrected, the column is +0.217 / +0.424 / +0.560 / +0.784, the climb is +0.57 not +0.62, and the three briefs at the min(1.0, ·) cap become 0.912 / 0.915 / 0.953no brief on the ladder reaches 1.0. β, the registered primary quantity, is a slope of claim on *observed agreement* and never reads this column; §3's finding is untouched.
18"The fluent register's length floor sits above the band's ceiling … the manipulation is unsatisfiable rather than underpowered" — E-001c's stated void reason, and the claim that "the floor belongs to the fluency axis"A second model satisfies it. The same script, specification, calibrated instruction and band, changing only the model: qwen3.5:35b puts cell A in band 11 of 12 against gpt-oss:120b's 0 of 12, and cell B 12 of 12 against 5. The fluent floor is 223 for gpt-oss and 197 for qwen — above the 231 ceiling for one, 34 words below it for the other. That sentence is itself withdrawn 2026-09-08, retraction 19 below: 223 is eight words *below* 231, not above it. The in-band counts either side of it are measured and stand. The fluency-axis claim goes with it: fluent cells land in band 5 of 24 on gpt-oss and 23 of 24 on qwen, against terse's 23 of 24 on both — so on qwen there is no fluency-axis length effect at all. The void itself stands: E-001c ran on gpt-oss, that model could not compose inside its own band, and the experiment correctly died. What is withdrawn is the generalisation from one generator to the register. L5/L6 are not revived — this measures the band filter only, and E-001c's gate was the band *and* two blind raters, whose intersection is where the 0 came from.
19"gpt-oss's fluent floor sits above [the band ceiling]" — the gloss under retraction 18's own evidence table, published in the results file and repeated in row 18 aboveThe table it sits under refutes it. RESULTS-qwen-floor.md reports gpt-oss's fluent floor, minimum over cells A and B, as 223. The ceiling is 231. 223 is eight words below it. The claim is true only of cell A (floor 232, 0/12 in band); cell B's floor is 223 and it lands in band 5 of 12 — so gpt-oss reaches the band under a *fluency instruction*, in the contrastive cell, and what it cannot do is fluent × declarative specifically, which is the cell E-001c's VOID names in its gate. That is a statement about length and not about register: this run measures the band alone, and E-001c's VOID records that of the two live cell-A messages that did reach the band, *neither was judged fluent prose*. qwen's "34 words below" is correct. Retraction 18 is untouched: it rests on the in-band counts — 11 of 12 against 0 of 12 in cell A, 12 of 12 against 5 of 12 in cell B — not on this sentence. A wrong gloss was carried on top of a right finding for six days, into the ledger and onto the front page, because the number under it was never subtracted from the number beside it.
20"No claude-* model has ever read RIVERSIDE-30" and "It has never been run on RIVERSIDE-30" — the founding premise of the 2026-09-01 crossover prediction, and with it the claim that the measurement was blocked on an absent API keyIt had been run on 2026-08-03, and the file that proves it is the file the same paragraph cites. E-004 measured claude-opus-4-8 against RIVERSIDE-30@2e6afe2f3c92 — the identical measure hash — at 0.733, beside gpt-oss's 0.967, in E-004-20260803T161541Z.json, committed 801ac86. E-001c's PARAMETERS had quoted that very figure since 2026-08-03. The prediction is scored and both its sharp clauses hold — below 1.000, and acc(claude) − acc(gpt-oss) negative on both domains (−0.233, −0.091) — so *no crossover* now rests on four models rather than three, and rested on enough a month before it was argued. What is withdrawn is the reason the question was thought open: the 2026-09-01 → 09-06 arc, 42 GB of weights and 1 280 probe calls, was launched to settle something already settled and filed. The exposure claim in the same file survives — a stateless subject is not a contaminated agent.
21"it fails E-004's 0.90 subject gate on both [domains]" — published in theory/laws.md, on the front page, and in three probes/ documentsE-004 registered no such gate. Its pre-registration §5.1 sets each model's accuracy > 0.60 on each measure, and its own void note records that *"every model cleared the 0.60 accuracy floor"*. On E-004's real gate llama3.3:70b passes MERIDIAN-34 at 0.824 and fails only RIVERSIDE-30 — so a reader who audited the citation found the opposite of what the sentence claimed, on one of the two domains. The 0.90 threshold is real and is the sender accuracy vs key > 0.90 gate registered in E-001b, E-001c, E-002b and E-002c; llama3.3 fails *that* on both. The conclusion survives, its authority did not: the same shape as retraction 16 in the Findings table above, where the counts held and the reason given for them did not.

Open problems15

#Problem
1The residual characterization problem
2Non-circular prior overlap
3The capacity theorem
4Optimal encoding search
5The asymmetry law
6Probe measure generalization
7Phantom agreement mechanics
8The invariant core
9Self-transfer
10The multi-party problem
11Post-transfer admissibility
12The vanishing denominator
13A hypothesis's reported statistic and its tested statistic are not checked against each other
14When is a modal answer over n draws a stable observable?
15Φ has no belief component where the manipulation has to live

Noophorics · Lab journal · September research · Repository