From brief to behaviour.
Did the model recover the decision boundaries in the brief—or only produce a fluent paraphrase?
Open research programme · v0.5
νόος nous, mind + φορά phora, carrying
Measure what survives a handoff.
Every handoff makes a claim: that another mind would make the same decisions after the transfer. Noophorics develops falsifiable measurements for testing that claim across people, models, and sessions.
Experimental by design. No law is established, and several of our own claims have already been refuted or revised.
Two minds are given the same problem. One explains it to the other. Both agree the explanation landed. Nobody measures what was lost.
Principia Noophorica §0 — the gap
A coding agent hands a migration task to another agent. Against a stated probe measure of ten held-out decisions, both predict strong agreement. The observed decisions tell a different story.
Φ = mean(90%, 80%) − 60% = 25 percentage points. Both agents are confident, yet their decisions diverge. Whether Φ is reliable and useful across domains remains an open empirical question.
What matters is not whether the artifact looks complete, but whether the receiver reconstructs the decisions that matter.
Did the model recover the decision boundaries in the brief—or only produce a fluent paraphrase?
Did the artifact preserve the constraints the next agent needs to act correctly?
What survived summarisation, and which decisions changed after the context was compressed?
Noophorics is a research programme—not a validated standard or a production product. Its current instruments and record are available for inspection and attack.
Divergence, transfer fidelity and its decomposition, with tests and synthetic validation.
Run the metrics →A schema for declaring the decisions against which a transfer is measured.
Inspect the probes →A handoff protocol carrying constraints, probes and a fidelity claim. v0.1 shipped as prose alone; implementing it found four defects — including a gate that passed a handoff where nothing transferred and both parties knew it. Repaired, and left visible.
Inspect NHP-0001 →Pre-registrations, void and null results, corrections and refuted claims remain visible.
Review the evidence →Information theory measures how much uncertainty a message removes — assuming sender and receiver share a codebook. Between a person and a model, between two models, or between a session and the summary it inherits, the codebooks differ, are private, and are partially reconstructed on the fly. Shannon's capacity is defined over symbols. What is needed is a capacity defined over reconstructed dispositions.
This happens millions of times a day, and every instance may be lossy. What is still missing is one instrument that reports how lossy, what was lost, and whether a better encoding existed while also measuring how well both parties calibrated their confidence.
The neighbouring fields each supply part of the apparatus and stop short:
| Field | What it gives | Where it stops |
|---|---|---|
| Information theoryShannon, 1948 | Uncertainty reduction, channel capacity, coding theorems. | Assumes a shared codebook. That assumption is exactly the problem. |
| PragmaticsGrice · Clark · Sperber & Wilson | Common ground, implicature, relevance — the right vocabulary. | Describes the phenomenon well and measures it barely. No units. |
| Rational Speech ActsFrank & Goodman, 2012 | Formal recursive listener–speaker reasoning. | A model of a dialogue, not a theory of the channel between two architectures. |
| Knowledge distillationteacher → student | Behaviour transferred between networks, measurably. | One-directional, requires a shared task, never asks what an optimal transfer would look like. |
| Alignmentdispositions we want | Whether a system's dispositions are the intended ones. | Noophorics asks a prior, more mechanical question: how much of a disposition arrives when we move it across a boundary? |
An earlier version of this page claimed there was no theory for this case. That was an overclaim, and a self-undermining one for a programme whose credibility rests on calibration. Adjacent literatures exist and are close: decision-preserving compression, semantic rate–distortion for heterogeneous agents, goal-oriented semantic communication, and learning a black-box receiver.
The defensible claim is narrower:
a unified measurement framework for decision-preserving transfer
between black-box agents, carrying fidelity, cost, and the calibration of
both parties in one instrument. The calibration term —
Φ — is the part we have not found
elsewhere.
Refuted 2026-07-29. The gap between believed and actual transfer has been measured, with numbers, since at least 1990 — Keysar & Henly elicited it per trial from 40 speaker–listener pairs in 2002, Chang et al. measured it in clinical handoff in 2010, Endsley reviewed 37 studies of the same divergence in 2020. Their instruments are in places better than ours. What survives is a claim about coverage, not discovery: one instrument carrying fidelity, cost and both parties’ calibration against a single stated probe measure. As for the gap itself between two language models — as of 2026-07-31 it is measured: E-002b reports Φ = +0.2961, CI [+0.2200, +0.3649], with one model in both roles, and that number is the parties’ near-constant 95.5% claim rate minus their 65.9% actual agreement. The level exists. Whether confidence responds to transfer is not measured and is registered as E-002c.
We refuse to define understanding as a state, because states are private and not comparable across architectures. We define it through behaviour:
B understands what A understands, with respect to a domain, to the extent that B would make the same decisions A would make, over that domain.Amended v0.4. B understands what A understands, with respect to a domain, to the extent that B's decisions over that domain move toward a stated reference for that domain. Where the reference is A's own decisions, what is measured is replication of A — a different and weaker claim, and the one this sentence had been making without saying so. Principia Noophorica §3
This is not a philosophical claim about what understanding is. It is a measurement convention — the same kind of move as defining temperature by the expansion of mercury rather than by the felt sensation of heat. It buys comparability across substrates, repeatability, falsifiability, and a number.
It also forces the field's most important structural discipline: there is no such thing as understanding in general — only understanding relative to a probe measure. A probe measure P is a distribution over decidable decisions. Fixing P is the noophoric equivalent of fixing a frame of reference in mechanics. Two people arguing about whether a model "really understood" the brief are, nine times out of ten, holding different P and not saying so.
Four axioms carry the weight. A1 — understanding is measured by divergence of decisions, not by symbol recovery, self-report, or surface similarity of text. A2 — every quantity is stated relative to a probe measure; a fidelity number without its P is as meaningless as a velocity without a frame. A3 — against probes the sender has never seen, and for messages of bounded cost, capacity is strictly less than 1; capacity is a curve K(C), not a scalar. In that restated form it is the quantitative version of Quine's indeterminacy of translation. The v0.1 form, which forgot both conditions, was refuted by a 113-token lookup table. A4 — the parties' confidence that a transfer succeeded is an independent observable, not a proxy for whether it did.
A4 is the axiom the rest of this page turns on. The difference between belief and measurement has a name, a symbol, and a section: §4.
The science of communication never got its instrument, because it could never open the receiver. You can ask a listener what they understood, but the report is not the state — people are unreliable narrators of their own comprehension, and you cannot run ten thousand controlled trials on one human mind.
| 1608 | The telescope | made astronomy. |
| 1670s | The microscope | made microbiology. |
| 1890s | The oscilloscope | made electronics an empirical discipline rather than a set of maxims. |
| now | The instrumentable mind | should make noophorics — or should fail to, publicly. |
Language models change the situation. A receiver can be resampled to build an empirical distribution over its dispositions, probed on arbitrary decisions, ablated at the input, and put through the whole procedure ten thousand times at a cost measured in cents. The same is true of the sender. That is the instrument. This programme is what we are trying to build with it.
The fourth row is a claim, not a record. The first three happened; the fourth is what this repository is betting on.
Divergence is Jensen–Shannon, in bits, estimated from n independent samples per probe. JSD is chosen because it is symmetric — noophoric divergence should not depend on which agent we call the sender — finite for distributions with disjoint support, and bounded in [0, 1] at base 2.
| Symbol | Name | Definition |
|---|---|---|
| P | Probe measurethe frame | A distribution over probes — decisions whose answer space is finite and discrete. No noophoric quantity is defined without one. A probe measure is admissible for a transfer only if the two agents actually disagree on it beforehand; measuring fidelity where agreement already exists measures nothing. |
| D | Divergence§2.2 | The expected Jensen–Shannon divergence between two agents' answer distributions over P. Bounded [0, 1]. D = 0 means behaviourally indistinguishable; D = 1 means they never give the same answer. |
| Â | Agreement rate§2.3 | The fraction of probes on which the two agents' modal answers match. Coarser than D, and used for phantom agreement because it is the quantity parties can actually estimate about themselves — nobody has calibrated intuitions about expected JSD. |
| Dfloor | Noise floor§3.2 | Finite-sample estimator bias at a stated n, obtained by a permutation null: pool both agents' draws for a probe, reshuffle into two groups of the original sizes, take the expected divergence. Under the null both groups come from one distribution, so this is what a perfectly aligned pair scores at that n — a property of the measurement, not of the agents. |
| F*R | Transfer fidelity§4.1 | The fraction of the pre-existing, closable gap toward a declared reference
R that the message actually closed. F* = 1 — the gap is closed, up to noise. F* = 0 — the message changed nothing. F* < 0 — the message made things worse: an antinoophor. Reported unclipped, because clipping hides antinoophors, which are among the most informative observations in the field.
v0.4: the reference is an argument, not an assumption. For three versions this measured movement toward the sender and never said so — and E-001 measured the cost: both receivers out-decided the sender against the key, fidelity ranked them the other way, and 62% of the effect sat on the four probes where the sender was wrong. Where R is the sender the quantity is replication, which is a real thing to measure and is not understanding. F*R=sender is identically the old number, so nothing published moved. |
| η | Efficiency§4.3 | F* / C, where cost C defaults to tokens measured with the receiver's tokenizer — the receiver is who pays to read it. Understanding per unit cost, valid only where F* ≥ 0. |
| Φ | Phantom agreement§5 · see below | Claimed agreement minus observed: Φ = Ĉ − Â, where Ĉ is the mean of the sender's prediction and the receiver's self-report. Both parties believe understanding transferred; probes say otherwise. The field's central pathology. |
| β | Calibration slope§5.1 · new in v0.5 | d(claimed) / d(observed) — how far a party's claim moves when the outcome moves. 1 is calibrated, 0 is inert, below zero is anti-calibrated. Independent of Φ, and reporting one without the other is the same error as reporting bias without resolution: a party that claims the long-run mean every time scores Φ = 0 and β = 0, and knows nothing. Measured in E-002c: sender −0.02, receiver +0.28, against 1 for calibration. |
| K | Channel capacity§6.1 | The best fidelity achievable by any message, at unbounded cost — the noophoric analogue of Shannon capacity and the central theoretical object of the field. Not directly computable; estimated as a lower bound over a stated search budget. We conjecture K < 1 always. |
| U | Residual§6.2 | 1 − K. The untransferable remainder. Axiom A3 asserts it is nonzero in the general case; characterising what lives inside it is Problem 1. |
Never report a fidelity without the floor correction, and never report one without naming the probe measure it was taken against. A number reported without its probe set, sample count, noise floor, cost unit, and agent identities is an anecdote. Anecdotes are accepted in the lab journal; they are not accepted in the theory.
Both parties believe understanding occurred. Probes reveal it did not. Principia Noophorica §5
Phantom agreement is to noophorics what dark matter is to
cosmology: the thing everyone had felt, nobody had weighed, and which we
suspect dominates the system. Refuted 2026-07-29: it has been
weighed in humans many times, and “dominates the system” is a claim about
magnitude we have no measurement to support. It is dangerous precisely because it is
invisible from inside. Neither party has any signal that anything went wrong.
The sender has discharged their intent; the receiver has a coherent,
confident, wrong model; and the error surfaces only downstream, in an action
nobody traces back to the conversation.
The sign of Φ is the whole reading:
| Φ > 0 | Shared illusion | Both parties overestimate how much landed. The expected case, and the one the field exists to catch. |
| Φ ≈ 0 | Calibrated | Confidence tracks measurement. If this turns out to be the norm, noophorics loses its motivating phenomenon. |
| Φ < 0 | Mutual underconfidence | More transferred than either party believes. Rarer, and its own kind of failure — it causes redundant re-explanation and unnecessary escalation. |
Report the sender's and receiver's claims separately as well as their mean. They are frequently asymmetric, and the asymmetry is data.
A law enters the record only with a refutation condition attached, and leaves it never — a refuted law is struck through and kept in place, because knowing what is false is the larger part of the record. All six are conjectured — and three of them, L2, L4 and L6, carry a headline form that has already been struck, shown below in place.
The more fluent and well-organised a message, the more Φ rises — and it rises faster than F* does.
Eloquence increases the belief that understanding transferred more than it increases the transfer. Fluency is a signal both parties read as comprehension: the sender feels discharged because the artifact is well-formed; the receiver feels informed because it is easy to process. Neither feeling is evidence about decisions. Processing fluency is a known source of misplaced confidence in human cognition, and we conjecture it is at least as strong in systems trained on human text — possibly stronger, since fluency is closer to their training objective than accuracy of transfer is.
This is an indictment of the systems writing most of today's handoffs, including the one that drafted the founding documents. That is a reason to test it, not a reason to soften it.
Refuted if: at equal cost, narrative and contrastive encodings produce statistically indistinguishable Φ, or narrative produces lower Φ.
At equal cost, encoding "the cases where we would diverge" transfers more fidelity than encoding "what I understand."
Do not send the model. Send the boundaries of the model. The receiver already has a prior; a declarative description spends cost re-encoding the parts of the sender's model the receiver would have reconstructed anyway. A contrastive encoding spends cost only on the delta — precisely the probes where the two agents currently disagree. Under the definition of F*, which measures gap closure rather than information delivered, the contrastive encoding is spending every token on the numerator.
This is the field's first engineering prescription, and it
directly contradicts how nearly every handoff, summary, and spec is written
today. Refuted 2026-07-29. Structured-handoff prescriptions
with mandated content slots and a receiver-synthesis step have been deployed and
outcome-measured since 2014. The I-PASS bundle, across nine pediatric residency
programmes and 10 740 admissions, was associated with a 23% reduction in
medical errors (24.5 → 18.8 per 100 admissions) and a 30% reduction in
preventable adverse events, both P < 0.001, without lengthening
handoffs (Starmer et al., NEJM 371(19), 2014). What L6 still claims
is the contrast — I-PASS was compared against unstandardised usual
care, not against declarative prose at equal cost — and it claims it for
machine receivers, about which I-PASS says nothing. Still testable; no longer a
primacy claim.
Refuted if: contrastive encodings show η ≤ declarative encodings at equal cost, across domains.
Together they predict something uncomfortable: the encoding that transfers best is the one that feels worst — and every party's subjective sense of a good handoff would be anticorrelated with its quality.
E-001 tests both jointly, and is pre-registered so that a null result is publishable and damaging.
Numbering is stable; solved problems are annotated, not renumbered. Grouped here by what a solution would settle. The ten below are the founding set. Five more have been raised since, every one from a defect found in this programme’s own measurements — 11 post-transfer admissibility, 12 the vanishing denominator, 13 a reported statistic and its tested statistic not checked against each other, 14 when a modal answer over n draws is a stable observable, and 15 Φ having no belief component where the manipulation has to live. They are stated in full in theory/open-problems.md; 14 and 15 are the two currently driving the work.
Axiom A3 asserts a nonzero residual. These three ask how big it is, what is inside it, and whether its size is derivable rather than merely measurable.
L6 names a family that should beat another family. These ask for the optimum, and for what survives being passed along.
One problem is the most deployable in the list; the other decides whether this is a science or a heap of incommensurable benchmarks.
All the definitions above are dyadic. Real systems are graphs, and one of the most economically weighted cases is an agent talking to itself.
Every experiment pre-registers its hypothesis in a committed file before any data exists. The git history is the pre-registration record. Null results are committed with the same prominence as positive ones.
Two messages of equal token cost, same sender, same source material, differing only in encoding style. Pre-registered 2026-07-28, amended three times on the instrument, and never completed — the sender model refused to compose the briefs at roughly nine attempts in ten, a refusal measured to be model-specific and unstable over time.
It nonetheless produced this programme’s only completed result, and the result is against the programme. Recomputed from the partial run: both receivers out-decided the sender that briefed them — accuracy 0.971 and 0.941 against the sender’s 0.882 — while fidelity ranked them the other way. On the four probes the sender got wrong, one receiver copied its errors and was rewarded; the other answered correctly and was charged near-maximal divergence. 62% of the headline effect came from those four probes.
So F* cannot separate reconstructing a domain from reconstructing the sender’s defects, and a transfer leaving the receiver more competent than the sender scores as partial failure. That is a defect in the definition, not the estimator: no floor, no sample size and no amendment repairs it. Fidelity is now decomposed into convergence-where-the-sender-is-right, error replication, and gap closure bought by class-prior matching alone — which in this run was worth 0.403 on its own.
The pre-registered test on the partial data returns p = 0.123. A null. Both L5 and L6 remain conjectured; E-001 tested neither.
The cost-parity gate failed on the composed messages: fluent briefs cost 1.5× terse ones under an identical budget instruction. Style and length are entangled in the generator.
The fluent register’s length floor sits above the word band’s
ceiling — 232 words against 231 — for gpt-oss:120b,
and for cell A only — corrected 2026-09-08
(retraction 19): 232 is the floor of the
fluent × declarative cell. The fluent register’s floor,
taken across both fluent cells, is 223 — eight words
below the ceiling, not above it.
Corrected 2026-08-31: the same script, specification, instruction and band on
qwen3.5:35b put cell A in band 11 of 12 against 0 of 12, and
showed no fluency-axis length effect at all — 23 of 24 fluent against 23 of 24 terse.
The void stands; “unsatisfiable rather than underpowered” was a generalisation
from one generator and is withdrawn.
330 of 330 elicitations returned “yes, we will agree” — the instrument had a default answer, because it ported Keysar & Henly’s granularity and not their forced-choice structure. The transfer was also perfect, 0 of 33 probes diverging, so there was nothing for a belief to be wrong about. This row said “Ablation ladder · Planned” until 2026-08-19, because the founding roadmap gave E-002 to the ladder and the identifier was later reused. The ladder is now E-006.
Four cost rungs, 24 briefs. Produced the first Φ measurements on this instrument and a headline pull-quote that was later withdrawn as retraction 12.
Is confidence responsive to transfer at all? Measured β = −0.02 for the sender and +0.28 for the receiver against 1 for calibration, and supplied the outcome-variation gate of three that every later probe measure is costed against.
Fidelity measured over a grid of sender/receiver pairs, in both
directions. Needs multiple model families to be interesting.
Blocked on the instrument, 2026-08-20: the asymmetry there is reason
to expect between the two local models is about 0.219.
MERIDIAN-IX32 resolves 0.638 — its 32 probes are
about nine independent prompt templates. Powered 2026-08-20, still not testable.
RIVERSIDE-30, measured directly for the first time, returns 10.00 diverged
of 30 and its divergence is probe-attributable, so pairing buys
×2.58 ×2.46 for gpt-oss and
×1.80 for qwen, an MDE of
0.128 0.135–0.183 depending on the
reader — corrected 2026-09-08, the single figure was one
reader's, and the 0.219 beside it is an effect
derived on MERIDIAN, so a threshold from one measure was standing next to
an effect from another. But power was never the
only blocker: L2's sharper form needs a pair where domain prior and general
capability point in opposite directions, and qwen is stronger or level
on both domains. Measuring only that an asymmetry exists is retraction 5.
Blocked on model selection. Measured 2026-09-06: a third
lineage was added — llama3.3:70b — and it is uniformly
weaker, 0.824 and 0.467
against 1.000 and 0.967 — each model at its own default with the
think field omitted, which is default-versus-default
and not a matched regime, since llama3.3 has no reasoning mode at all. The
sign does not flip, so there is no crossover to find. It fails
E-004's 0.90 subject gate the > 0.90
sender-accuracy-vs-key gate registered in E-001b, E-001c, E-002b and E-002c
— not E-004's, which is 0.60 and which
llama3.3 passes on MERIDIAN-34
(retraction 21). And it was not the
first model asked. E-004 measured claude-opus-4-8 on 2026-08-03
against the identical RIVERSIDE-30 hash at 0.733
against gpt-oss's 0.967, and
0.909 against 1.000 on
MERIDIAN-33 — negative on both, no crossover, a month before the
question was posed as open
(retraction 20).
Blocked on domain selection — both domains are the same task
type, which is a probe measure and a month, not a download. The instrument
itself held up on a second reader (2026-08-28): qwen diverges 9.67 of 30 on
RIVERSIDE-30 and lands on the same probes, Jaccard
0.654 against a ≥ 0.4
predicted beforehand.
Whether a model can predict where another model will disagree, without a key. Void: two of the three registered models were never wrong — 1 error and 0 errors across the whole design — and a detector of disagreement needs disagreement to detect. This row said “Chain decay · Planned” until 2026-08-19, the same identifier reuse as E-002 above. The chain is now E-007.
Void before collection, 2026-08-20, at zero model calls. One message truncated at eight cost levels, with the resulting F*(C) curve fitted. Its own pre-registered power gate refused the run: power 0.117 against a moderately concave truth on 32 probes. L1's concavity limb needs about 512 probes; the largest measure this programme owns has 34. Not that L1 has not been tried — that no probe measure here can carry the test. Was E-002 in the founding roadmap; renumbered 2026-08-19 after that identifier was reused.
A telephone chain of six hops, content tagged by form, per-hop fidelity measured against the origin — to test whether constraints outlast descriptions. Was E-004 in the founding roadmap; renumbered for the same reason. Blocked on the instrument, 2026-08-20: L4c is a difference between two classes of probe, so the split halves each n — at most ~16 per class, resolving 0.479.
NHP-0001 is a draft handoff format that carries a fidelity claim. Its unsolved problem is named in the draft itself: the protocol asks the sender to write its own exam.
The first live run of E-001 was recorded as void rather than discarded, and produced a finding that outlives the experiment: a measurement instrument can be blocked by the safety behaviour of the system it measures, and the block can be silent, asymmetric, and shaped like a result. Any evaluation that varies prompt style across conditions is also varying classifier surface across conditions. Refusals must be treated as missing data, never as a zero.
Every one of them, with what killed it → A count nobody can audit is worse than no count.
Six conjectural laws, fifteen open problems, four completed runs — and one pre-registered finding. Two produced only corrections to our own instrument, one was voided by its own gate after it had collected everything, and the finding is E-002c. Every other claim in this programme should be read as a bet, and its confidence calibrated accordingly.
A science that cannot lose is not a science. Noophorics is wrong if:
The method is the only part of this that is not negotiable. Every experiment pre-registers before data exists. Every number carries its probe measure, sample count, and noise floor. Refuted laws are struck through, never deleted. Where a bound is claimed, an adversarial attempt to violate it is expected in the same commit.
The name takes the noo- root from νόος and disclaims any inheritance from Vernadsky's or Teilhard's noosphere. Nothing in this programme is mystical. Everything in it is supposed to have a number attached, or be deleted.
One item. It is here because the sentence above it used to read “nothing is established”, and that stopped being true on 2026-08-02.
The programme's only completed results are negative and about itself. Each is recorded where it was made, with the counterexample, rather than edited away:
The correct response to this is not agreement. It is a probe.