Noophorics.org
Open research programme Version 0.3 An instrument for measuring what a receiver can still decide after a transfer.

Open research programme  ·  v0.3

Noophorics

νόος nous, mind  +  φορά phora, carrying

Measure what survives a handoff.

Every handoff makes a claim: that another mind would make the same decisions after the transfer. Noophorics develops falsifiable measurements for testing that claim across people, models, and sessions.

Experimental by design. No law is established, and several of our own claims have already been refuted or revised.

Two minds are given the same problem. One explains it to the other. Both agree the explanation landed. Nobody measures what was lost.

Principia Noophorica §0 — the gap

Φ · 60 seconds The pathology Both sides think understanding transferred. Held-out decisions say otherwise.

Both sides think it landed. The probes disagree.

A coding agent hands a migration task to another agent. Against a stated probe measure of ten held-out decisions, both predict strong agreement. The observed decisions tell a different story.

Three boundaries One question Did the receiver reconstruct the decisions that matter, rather than merely the words?

One question, three transfer boundaries.

What matters is not whether the artifact looks complete, but whether the receiver reconstructs the decisions that matter.

01 · Human → model

From brief to behaviour.

Did the model recover the decision boundaries in the brief—or only produce a fluent paraphrase?

02 · Model → model

Across an agent handoff.

Did the artifact preserve the constraints the next agent needs to act correctly?

03 · Session → summary

Through context compaction.

What survived summarisation, and which decisions changed after the context was compressed?

Available today Public work Open, reproducible, unfinished—and published early enough to fail in public.

Open now. Unfinished by design.

Noophorics is a research programme—not a validated standard or a production product. Its current instruments and record are available for inspection and attack.

Implemented · tested

Reference metrics

Divergence, transfer fidelity and its decomposition, with tests and synthetic validation.

Run the metrics →
Specified · evolving

Probe format

A schema for declaring the decisions against which a transfer is measured.

Inspect the probes →
v0.2 · implemented, unvalidated

NHP-0001

A handoff protocol carrying constraints, probes and a fidelity claim. v0.1 shipped as prose alone; implementing it found four defects — including a gate that passed a handoff where nothing transferred and both parties knew it. Repaired, and left visible.

Inspect NHP-0001 →
Public · append-only

Research record

Pre-registrations, void and null results, corrections and refuted claims remain visible.

Review the evidence →
Version
0.3
Conjectural laws
6 (2 restated after refutation)
Open problems
10
Experiments run to completion
0
Own claims refuted
3
Confirmed findings
0
Last revised
Licence
Apache-2.0 · CC BY 4.0
§0 The gap Principia Noophorica §1 — why the existing sciences do not cover this.

Shannon assumed a shared codebook. Here the codebooks differ.

Information theory measures how much uncertainty a message removes — assuming sender and receiver share a codebook. Between a person and a model, between two models, or between a session and the summary it inherits, the codebooks differ, are private, and are partially reconstructed on the fly. Shannon's capacity is defined over symbols. What is needed is a capacity defined over reconstructed dispositions.

This happens millions of times a day, and every instance may be lossy. What is still missing is one instrument that reports how lossy, what was lost, and whether a better encoding existed while also measuring how well both parties calibrated their confidence.

The neighbouring fields each supply part of the apparatus and stop short:

Tab. 1 — Prior art, and what each one leaves undone
FieldWhat it givesWhere it stops
Information theoryShannon, 1948 Uncertainty reduction, channel capacity, coding theorems. Assumes a shared codebook. That assumption is exactly the problem.
PragmaticsGrice · Clark · Sperber & Wilson Common ground, implicature, relevance — the right vocabulary. Describes the phenomenon well and measures it barely. No units.
Rational Speech ActsFrank & Goodman, 2012 Formal recursive listener–speaker reasoning. A model of a dialogue, not a theory of the channel between two architectures.
Knowledge distillationteacher → student Behaviour transferred between networks, measurably. One-directional, requires a shared task, never asks what an optimal transfer would look like.
Alignmentdispositions we want Whether a system's dispositions are the intended ones. Noophorics asks a prior, more mechanical question: how much of a disposition arrives when we move it across a boundary?

An earlier version of this page claimed there was no theory for this case. That was an overclaim, and a self-undermining one for a programme whose credibility rests on calibration. Adjacent literatures exist and are close: decision-preserving compression, semantic rate–distortion for heterogeneous agents, goal-oriented semantic communication, and learning a black-box receiver.

The defensible claim is narrower: a unified measurement framework for decision-preserving transfer between black-box agents, carrying fidelity, cost, and the calibration of both parties in one instrument. The calibration term — Φis the part we have not found elsewhere.

Refuted 2026-07-29. The gap between believed and actual transfer has been measured, with numbers, since at least 1990 — Keysar & Henly elicited it per trial from 40 speaker–listener pairs in 2002, Chang et al. measured it in clinical handoff in 2010, Endsley reviewed 37 studies of the same divergence in 2020. Their instruments are in places better than ours. What survives is a claim about coverage, not discovery: one instrument carrying fidelity, cost and both parties’ calibration against a single stated probe measure — and no measurement of this gap anywhere, including here, where sender and receiver are both language models.

§1 The move Understanding defined operationally, the way temperature was defined by the expansion of mercury.

Understanding is measured by behavioural convergence.

We refuse to define understanding as a state, because states are private and not comparable across architectures. We define it through behaviour:

B understands what A understands, with respect to a domain, to the extent that B would make the same decisions A would make, over that domain. Principia Noophorica §3

This is not a philosophical claim about what understanding is. It is a measurement convention — the same kind of move as defining temperature by the expansion of mercury rather than by the felt sensation of heat. It buys comparability across substrates, repeatability, falsifiability, and a number.

It also forces the field's most important structural discipline: there is no such thing as understanding in general — only understanding relative to a probe measure. A probe measure P is a distribution over decidable decisions. Fixing P is the noophoric equivalent of fixing a frame of reference in mechanics. Two people arguing about whether a model "really understood" the brief are, nine times out of ten, holding different P and not saying so.

A SENDER B|m RECEIVER m THE ARTIFACT — COST C(m) D( A , B|m  |  P ) P — PROBE MEASURE · THE FRAME OF REFERENCE
Fig. 1 The instrument. Both agents are resampled against the same probe measure; what is measured is the divergence between their answer distributions, before and after the artifact. Nothing about either agent's internals is required — an agent is defined entirely by its answer distributions over a probe space.

Four axioms carry the weight. A1 — understanding is measured by divergence of decisions, not by symbol recovery, self-report, or surface similarity of text. A2 — every quantity is stated relative to a probe measure; a fidelity number without its P is as meaningless as a velocity without a frame. A3 — between any two systems with non-identical priors there exist dispositions that cannot be transferred at any message length; this is the quantitative form of Quine's indeterminacy of translation. A4 — the parties' confidence that a transfer succeeded is an independent observable, not a proxy for whether it did.

A4 is the axiom the rest of this page turns on. The difference between belief and measurement has a name, a symbol, and a section: §4.

§2 Why now Every science begins with an instrument.

The receiver is instrumentable for the first time.

The science of communication never got its instrument, because it could never open the receiver. You can ask a listener what they understood, but the report is not the state — people are unreliable narrators of their own comprehension, and you cannot run ten thousand controlled trials on one human mind.

Tab. 2 — Instrument, then field
1608 The telescope made astronomy.
1670s The microscope made microbiology.
1890s The oscilloscope made electronics an empirical discipline rather than a set of maxims.
now The instrumentable mind should make noophorics — or should fail to, publicly.

Language models change the situation. A receiver can be resampled to build an empirical distribution over its dispositions, probed on arbitrary decisions, ablated at the input, and put through the whole procedure ten thousand times at a cost measured in cents. The same is true of the sender. That is the instrument. This programme is what we are trying to build with it.

The fourth row is a claim, not a record. The first three happened; the fourth is what this repository is betting on.

§3 The quantities theory/definitions.md — formal definitions; metrics/ — reference implementation, stdlib only.

Nine quantities, each with a definition you can implement.

Divergence is Jensen–Shannon, in bits, estimated from n independent samples per probe. JSD is chosen because it is symmetric — noophoric divergence should not depend on which agent we call the sender — finite for distributions with disjoint support, and bounded in [0, 1] at base 2.

Tab. 3 — Core quantities
SymbolNameDefinition
P Probe measurethe frame A distribution over probes — decisions whose answer space is finite and discrete. No noophoric quantity is defined without one. A probe measure is admissible for a transfer only if the two agents actually disagree on it beforehand; measuring fidelity where agreement already exists measures nothing.
D Divergence§2.2 The expected Jensen–Shannon divergence between two agents' answer distributions over P. Bounded [0, 1]. D = 0 means behaviourally indistinguishable; D = 1 means they never give the same answer.
 Agreement rate§2.3 The fraction of probes on which the two agents' modal answers match. Coarser than D, and used for phantom agreement because it is the quantity parties can actually estimate about themselves — nobody has calibrated intuitions about expected JSD.
Dfloor Noise floor§3.2 The mean self-divergence of the two agents: the divergence that two perfectly aligned agents would still show, from their own stochasticity. No transfer can push measured divergence below it. Not correcting for it is the field's most common error — an uncorrected fidelity number silently attributes sampling noise to untransferred meaning.
F* Transfer fidelity§4.1 The fraction of the pre-existing, closable gap that the message actually closed. F* = 1 — the gap is closed, up to noise. F* = 0 — the message changed nothing. F* < 0 — the message made things worse: an antinoophor. Reported unclipped, because clipping hides antinoophors, which are among the most informative observations in the field.
η Efficiency§4.3 F* / C, where cost C defaults to tokens measured with the receiver's tokenizer — the receiver is who pays to read it. Understanding per unit cost: the quantity engineering should optimise, and almost nothing currently does.
Φ Phantom agreement§5 · see below Claimed agreement minus observed: Φ = Ĉ − Â, where Ĉ is the mean of the sender's prediction and the receiver's self-report. Both parties believe understanding transferred; probes say otherwise. The field's central pathology.
K Channel capacity§6.1 The best fidelity achievable by any message, at unbounded cost — the noophoric analogue of Shannon capacity and the central theoretical object of the field. Not directly computable; estimated as a lower bound over a stated search budget. We conjecture K < 1 always.
R Residual§6.2 1 − K. The untransferable remainder. Axiom A3 asserts it is nonzero in the general case; characterising what lives inside it is Problem 1.
D(A, B | P) − D(A, B|m | P) F*(m) = ───────────────────────────────── D(A, B | P) − D_floor(A, B | P)

Never report a fidelity without the floor correction, and never report one without naming the probe measure it was taken against. A number reported without its probe set, sample count, noise floor, cost unit, and agent identities is an anecdote. Anecdotes are accepted in the lab journal; they are not accepted in the theory.

§4 The pathology Axiom A4 — belief is not evidence. The difference is itself a quantity.

Φ — phantom agreement

Both parties believe understanding occurred. Probes reveal it did not. Principia Noophorica §5

Phantom agreement is to noophorics what dark matter is to cosmology: the thing everyone had felt, nobody had weighed, and which we suspect dominates the system. Refuted 2026-07-29: it has been weighed in humans many times, and “dominates the system” is a claim about magnitude we have no measurement to support. It is dangerous precisely because it is invisible from inside. Neither party has any signal that anything went wrong. The sender has discharged their intent; the receiver has a coherent, confident, wrong model; and the error surfaces only downstream, in an action nobody traces back to the conversation.

Φ = Ĉ − Â Ĉ claimed agreement — the mean of the sender's prediction and the receiver's self-report, each elicited as: "over probes of this kind, what fraction of your decisions would match the other party's?" Â observed agreement — the fraction of probes on which the two agents' modal answers actually match.
1.0 0 R — RESIDUAL (1 − K) K — CAPACITY Φ Ĉ — CLAIMED Â — OBSERVED MESSAGE COST C(m) → AGREEMENT RATE
Fig. 2 Schematic — not data. Claimed agreement outruns observed agreement; the area between them is Φ. Observed agreement saturates at the capacity K, leaving the residual R = 1 − K. No curve on this plate has been measured. It is what axiom A3 and law L5 predict, drawn so that it can be looked at and disagreed with.

The sign of Φ is the whole reading:

Φ > 0 Shared illusion Both parties overestimate how much landed. The expected case, and the one the field exists to catch.
Φ ≈ 0 Calibrated Confidence tracks measurement. If this turns out to be the norm, noophorics loses its motivating phenomenon.
Φ < 0 Mutual underconfidence More transferred than either party believes. Rarer, and its own kind of failure — it causes redundant re-explanation and unnecessary escalation.

Report the sender's and receiver's claims separately as well as their mean. They are frequently asymmetric, and the asymmetry is data.

§5 The laws theory/laws.md — six conjectures, each with a stated kill condition. None is established.

Six conjectures. Two of them are the argument.

A law enters the record only with a refutation condition attached, and leaves it never — a refuted law is struck through and kept in place, because knowing what is false is the larger part of the record. All six currently hold the same status: conjectured

L1 Diminishing noophoric return. F*(C) is concave in message cost and saturates strictly below 1. E-002
L2 Asymmetry. F*(A→B) ≠ F*(B→A), with the sign predictable from the asymmetry of the two priors — and we expect domain prior to dominate general capability. E-003
L3 Prior overlap. K rises monotonically with the overlap of the agents' priors, and below a minimum shared basis no message length helps. Needs Problem 2
L4 The curse of the summary. Fidelity is multiplicative along a chain — but an invariant core survives arbitrarily many hops, and we conjecture that core is constraints and prohibitions rather than descriptions. E-004
L5 Fluency inflates phantom agreement. See below. E-001
L6 Optimal encoding is contrastive. See below. E-001
L5

Fluency inflates phantom agreement

Conjectured

The more fluent and well-organised a message, the more Φ rises — and it rises faster than F* does.

Eloquence increases the belief that understanding transferred more than it increases the transfer. Fluency is a signal both parties read as comprehension: the sender feels discharged because the artifact is well-formed; the receiver feels informed because it is easy to process. Neither feeling is evidence about decisions. Processing fluency is a known source of misplaced confidence in human cognition, and we conjecture it is at least as strong in systems trained on human text — possibly stronger, since fluency is closer to their training objective than accuracy of transfer is.

This is an indictment of the systems writing most of today's handoffs, including the one that drafted the founding documents. That is a reason to test it, not a reason to soften it.

Refuted if: at equal cost, narrative and contrastive encodings produce statistically indistinguishable Φ, or narrative produces lower Φ.

L6

Optimal encoding is contrastive, not declarative

Conjectured

At equal cost, encoding "the cases where we would diverge" transfers more fidelity than encoding "what I understand."

Do not send the model. Send the boundaries of the model. The receiver already has a prior; a declarative description spends cost re-encoding the parts of the sender's model the receiver would have reconstructed anyway. A contrastive encoding spends cost only on the delta — precisely the probes where the two agents currently disagree. Under the definition of F*, which measures gap closure rather than information delivered, the contrastive encoding is spending every token on the numerator.

This is the field's first engineering prescription, and it directly contradicts how nearly every handoff, summary, and spec is written today.

Refuted if: contrastive encodings show η ≤ declarative encodings at equal cost, across domains.

Together they predict something uncomfortable: the encoding that transfers best is the one that feels worst — and every party's subjective sense of a good handoff would be anticorrelated with its quality.

E-001 tests both jointly, and is pre-registered so that a null result is publishable and damaging.

§6 Roadmap theory/open-problems.md — stated in Hilbert's spirit: precise enough to work on, open enough to be hard.

Ten problems for the first decade.

Numbering is stable; solved problems are annotated, not renumbered. Grouped here by what a solution would settle.

I — The ceiling · what cannot be transferred

Axiom A3 asserts a nonzero residual. These three ask how big it is, what is inside it, and whether its size is derivable rather than merely measurable.

01The residual characterisation problem. What lives inside R = 1 − K? Given two agents and a domain, predict which dispositions fall inside the residual before measuring.
02Non-circular prior overlap. Measure the overlap of two agents' priors without using transfer fidelity to do it. Until this is solved, L3 is a tautology — and it is currently the biggest hole in the theory.
03The capacity theorem. Is there a Shannon-style coding theorem for mismatched priors — does K admit a closed form, or is it irreducibly empirical? The field's central theoretical question. A negative answer is also a result.

II — The encoding · what to send

L6 names a family that should beat another family. These ask for the optimum, and for what survives being passed along.

04Optimal encoding search. Given (A, B, P, C_max), construct the message maximising F*. Practical payoff is immediate: this is the algorithm every agent handoff should be running.
08The invariant core. Prove or refute that constraints survive chained transfer where descriptions do not, and characterise the invariant class exactly. Candidate axis: dispositions whose verification is cheaper than their derivation.

III — The pathology · what goes wrong, and whether we can trust the measurement

One problem is the most deployable in the list; the other decides whether this is a science or a heap of incommensurable benchmarks.

07Phantom agreement mechanics. What generates Φ, and what is the cheapest reliable detector? Find the minimal probe set that detects Φ > θ at a stated sensitivity, cheap enough to fire on every production handoff.
06Probe measure generalisation. When do fidelity results on P₁ predict fidelity on P₂? This is also the field's overfitting problem — without it, solutions to Problem 4 cannot be trusted.

IV — The topology · who transfers to whom

All the definitions above are dyadic. Real systems are graphs, and one of the most economically weighted cases is an agent talking to itself.

05The asymmetry law. Predict the sign and magnitude of F*(A→B) − F*(B→A) from properties of A and B. In particular, resolve the competition between general capability and domain prior.
09Self-transfer. How much of an agent survives its own compaction? Priors match perfectly, but the compaction artifact is generated under the same fluency pressure as any other message — so L5 predicts self-transfer should exhibit maximal Φ.
10The multi-party problem. Define fidelity over a topology, then find the communication graph maximising end-to-end fidelity under a cost budget. This turns multi-agent architecture from a design metaphor into an optimisation problem.

Experiments

Every experiment pre-registers its hypothesis in a committed file before any data exists. The git history is the pre-registration record. Null results are committed with the same prominence as positive ones.

E-001 The cost of fluency attacks L5, L6 Void · informative

Two messages of equal token cost, same sender, same source material, differing only in encoding style. Pre-registered 2026-07-28, amended three times on the instrument, and never completed — the sender model refused to compose the briefs at roughly nine attempts in ten, a refusal measured to be model-specific and unstable over time.

It nonetheless produced this programme’s only completed result, and the result is against the programme. Recomputed from the partial run: both receivers out-decided the sender that briefed them — accuracy 0.971 and 0.941 against the sender’s 0.882 — while fidelity ranked them the other way. On the four probes the sender got wrong, one receiver copied its errors and was rewarded; the other answered correctly and was charged near-maximal divergence. 62% of the headline effect came from those four probes.

So F* cannot separate reconstructing a domain from reconstructing the sender’s defects, and a transfer leaving the receiver more competent than the sender scores as partial failure. That is a defect in the definition, not the estimator: no floor, no sample size and no amendment repairs it. Fidelity is now decomposed into convergence-where-the-sender-is-right, error replication, and gap closure bought by class-prior matching alone — which in this run was worth 0.403 on its own.

The pre-registered test on the partial data returns p = 0.123. A null. Both L5 and L6 remain conjectured; E-001 tested neither.

E-002 Ablation ladder attacks L1 Planned

One message truncated at eight cost levels, with the resulting F*(C) curve fitted. The cheapest remaining experiment.

E-003 Asymmetry matrix attacks L2 Planned

Fidelity measured over a grid of sender/receiver pairs, in both directions. Needs multiple model families to be interesting.

E-004 Chain decay attacks L4 Planned

A telephone chain of six hops, content tagged by form, per-hop fidelity measured — to test whether constraints outlast descriptions.

E-005+ Protocol validation NHP-0001 Unscheduled

NHP-0001 is a draft handoff format that carries a fidelity claim. Its unsolved problem is named in the draft itself: the protocol asks the sender to write its own exam.

The first live run of E-001 was recorded as void rather than discarded, and produced a finding that outlives the experiment: a measurement instrument can be blocked by the safety behaviour of the system it measures, and the block can be silent, asymmetric, and shaped like a result. Any evaluation that varies prompt style across conditions is also varying classifier surface across conditions. Refusals must be treated as missing data, never as a zero.

§7 Standing Principia Noophorica §7 and §10 — what would falsify the programme.

Version 0.3. Nothing is established, and three of our own claims are refuted.

One pre-registered experiment that has not produced data, six conjectural laws, ten open problems, zero confirmed findings. Every claim in this programme should be read as a bet, and its confidence calibrated accordingly.

A science that cannot lose is not a science. Noophorics is wrong if:

  1. 1F* is not stable under probe resampling. If independently drawn probe measures over the same nominal domain give uncorrelated fidelity scores, the quantity is noise. — the most dangerous of the four.
  2. 2Φ is consistently ≈ 0. If the parties' confidence tracks measured fidelity, the central pathology does not exist and the field loses its motivating phenomenon.
  3. 3K ≈ 1 in practice. If sufficiently long messages reliably close the gap between arbitrary systems, axiom A3 is false and the interesting structure disappears.
  4. 4No encoding beats any other at equal cost. If η is invariant to message form, there is nothing to engineer and the field is descriptive at best.

The method is the only part of this that is not negotiable. Every experiment pre-registers before data exists. Every number carries its probe measure, sample count, and noise floor. Refuted laws are struck through, never deleted. Where a bound is claimed, an adversarial attempt to violate it is expected in the same commit.

The name takes the noo- root from νόος and disclaims any inheritance from Vernadsky's or Teilhard's noosphere. Nothing in this programme is mystical. Everything in it is supposed to have a number attached, or be deleted.

What we have refuted, in our own work

The programme's only completed results are negative and about itself. Each is recorded where it was made, with the counterexample, rather than edited away:

  • F* rewards replicating the sender's errors. Both receivers out-decided the sender; fidelity ranked them the other way. FINDINGS
  • Axiom A3 was false as first stated. A 113-token lookup table over a visible probe measure reaches F* = 1. PRINCIPIA §4
  • The v0.1 noise floor was inflated ~51%. Found by validating the estimator against a known ground truth — a check that did not exist until it was needed. validation
  • L2 was a tautology and L4 was ill-typed. Two antinoophors composed to a positive fidelity. laws
  • Our own first live run was void, and the near miss is the most useful thing in the notebook. lab journal

The correct response to this is not agreement. It is a probe.