instrument · qwen register feasibility

Qwen register feasibility: instrument readout

2026-09-14 (UTC start date). Collection complete.

The complete two-judge corpus result is available. Recorded HTTP rating attempts: 104 / 104 (fixed cap); invalid or indeterminate attempted responses: 0. Corpus messages missing at least one valid rating: 0 / 48.

Direct joint acceptance across the archived corpus: 32 / 48. The judges disagreed on 16 / 48 messages.

These are descriptive counts for this finite archived corpus and these configured judges. This readout introduces no scientific hypothesis, law, significance test, or population acceptance estimate.

Calibration

JudgeStatusControls attempted / plannedValid responsesCorrect controls
gpt-oss:120bpassed4 / 444
llama3.3:70bpassed4 / 444

Every control must be correct before that judge's corpus calls are eligible. A failed or missing control excludes its corpus ratings; it does not authorize a replacement judge or retry. These procedural controls are not a validated benchmark, and passing them does not establish correctness on borderline passages.

Corpus counts

CellExpected registerMessagesIn word bandgpt-oss:120b expectedllama3.3:70b expectedBoth expectedDisagreementsDirect joint acceptance
Aconnected prose1211712757
Bconnected prose1212812848
Clist1212121212012
Dlist1211125575

The word band is 182–231 whitespace-separated words. Direct joint acceptance counts the same messages passing the band and both judges' expected register; it is the observed intersection, not a product or minimum of marginal counts. Expected-register counts use valid observed ratings. Both-expected and disagreement counts use only messages with two valid ratings. A complete joint count is unavailable for a cell with any missing rating, or if either judge fails calibration.

Missing data

Cellgpt-oss:120b validgpt-oss:120b missingllama3.3:70b validllama3.3:70b missingTwo validAny missing
A120120120
B120120120
C120120120
D120120120

Missing includes unattempted, skipped, invalid, and indeterminate corpus ratings. An invalid response never becomes a list verdict. A reserved call interrupted before a definitive response stays missing and is never resent.

Scope and provenance

All 48 archived compositions were fixed before the register outcomes were collected; their length outcomes already existed. Every composition is included, including those outside the word band. There is one draw per eligible message and judge, with no retries or replacement compositions.

The configured regimes differ: gpt-oss:120b uses think=medium; llama3.3:70b omits the thinking field. Both use temperature 0.7. These are observations from these particular judges, not a controlled comparison of model ability. Different model lineages do not establish independent errors. The historical qwen3.5:35b composer digest is unknown.

Register acceptance does not test factual accuracy, contrastiveness, transfer, fidelity, phantom agreement, or usefulness. E-001c remains void, and its pre-registration is unchanged. Scientific interpretation requires human review; this file is a mechanical instrument readout.

Collection started: 2026-09-14T07:42:05.627766+00:00. Plan committed before ratings: 544e6fe. Frozen runner commit: d0de1a3.

Pinned inputSHA-256
Archived source JSONcfbd5d776cec1bb430f0bb9ec82b5d1a08f2c0a22c5dff43520a7f1fb8f5d0f9
Plan JSON8709fb5c79ed1c7d306ef15dac4a42d2e860fdd34359fd0f8126922d5cdb4f6c
Runner32a5120cb6c410fff2f0c9737fc71c353fd2348b4240d5722dabd954900757fe
Absolute question44571477c1c92e7ce9526c707a88fc08eda3075fa87459871153f1f7cdcb6590
gpt-oss:120b weightsa951a23b46a1f6093dafee2ea481d634b4e31ac720a8a16f3f91e04f5a40ecd9
llama3.3:70b weightsa6eb4748fd2990ad2952b2335a95a7f952d1a06119a0aa6a2df6cd052a93a3fa

Sources: prospective plan, machine-readable plan, archived source, result manifest and item identities, saved summary, raw call records.

Every count above is recomputed from checked raw call records, not copied from the saved summary. The manifest binds item identities and text hashes; the runner checks them alongside the source, question, plan, model, and runner identities.

Recompute from the saved raw records, starting at the repository root:

python3 tools/rate_register.py --out probes/qwen-register/results/20260914T074205Z --summarize
python3 tools/report_register.py --out probes/qwen-register/results/20260914T074205Z --report research/2026-09-14-register-feasibility.md --check

This document is licensed CC BY 4.0.

All journal entries · Reference · Noophorics · Repository