CR study — handoff to the main PC (weak rungs + final trend test)
Written from the Copilot PC after measuring the mid + frontier rungs. This doc tells the main-PC operator exactly what remains to complete the CR (capacity-relative) study and produce the pre-registered result. Read alongside
PREREGISTRATION.md(the design is authoritative) andRESULTS-copilot-arm-mid-frontier.md(what this arm found).
1. Current state
- Done (Copilot PC): all 4 hosted rungs —
gpt-mid,gpt-frontier,gemini-frontier,claude-frontier— generated (naive + gs, k=3) and measured (static + oracle). Results inRESULTS-copilot-arm-mid-frontier.md+cr_analysis_mid_frontier.json. Headline: on the behavioral oracle, naive ≥ GS at every measured rung (GS ties only at the claude frontier). - Not done (this doc): the two weak local rungs —
qwen25coder7b(rung 0) andqwen25coder32b(rung 1) — via Ollama. Perladder.jsonthese are the a-priori weakest and are where H1 predicts GS’s advantage is largest. Without them the monotone-trend test in PREREGISTRATION §7 cannot run. This is the load-bearing remaining work.
2. What to generate on the main PC
For each slug ∈ {qwen25coder7b, qwen25coder32b}, each cond ∈ {naive, gs}, each rep ∈ 0..2 (k=3): follow runner/COPILOT_GENERATION_PROMPT.md verbatim. The non-negotiable rules:
- Fresh session per generation. No memory across reps/conditions/models.
- The generator NEVER sees
benchmark/oracle/. It is the grader; showing it = teaching to the test and invalidates the cell. - naive =
benchmark/DOMAIN_SPEC.md+benchmark/naive/README.md+ the naive prompts. - gs = the
benchmark/gs/cascade (CLAUDE.mdfirst), then the gs prompts. Do NOT give the gs arm the rawDOMAIN_SPEC.md— the cascade is its spec (that is the experiment). - Stage each build at
runner/runs/<slug>__<cond>/<rep>/project/(this dir is gitignored). - Canary first, per model (cold recall): run
runner/canary_probe.mdagainst each qwen model from its name alone; near-zero recall confirms non-contamination. Log it (canary_results.jsonalready holds the hosted-model canary). - Ollama ran out of VRAM twice in AX2 (PREREGISTRATION §8): run the weak rung serially, one generation at a time, kill node zombies between runs. If 32B won’t fit, drop rung 1 — the 3-rung minimum (qwen7b → gpt-mid → frontier) still spans weak→mid→strong.
3. How to measure (same instrument, now cross-platform)
Run the two instruments directly with node (they are the designed entry points; measure_cr.sh is optional and only wraps them):
cd experiments/cr/runner
node static_cr.cjs # → static_cr.json (append-only; delete the file to re-measure fresh)
node conformance_cr.cjs # → conformance_cr.json
Both auto-discover cells under runs/, so once the qwen cells are staged they are picked up automatically. Both JSON outputs are gitignored (per-machine); commit only curated results.
Portability notes (the instrument was hardened on the Copilot/Windows PC):
- Hurl path:
conformance_cr.cjsuseshurlfrom PATH on non-Windows; on Windows it defaults to the 8.3 short path. Override anywhere withCR_HURL=/path/to/hurl. - DB:
ensureDb()auto-manages a dockerpastura-measurepostgres:16-alpine on port 5545; the app is served on PORT 4147. Needs docker + a Node withnpx.tsxshould be globally available (npm i -g tsx) — some cells declare ts-node, andserve()usesnpx tsx. - Disk: each cell’s
npm installis ~300–400 MB. Cleannode_modulesunderruns/between rungs if space is tight (find runs -name node_modules -type d -prune -exec rm -rf {} +). - Prisma: cells that omit the
prismadevDep are pinned to the installed@prisma/clientversion to dodge the brokenprisma@8.0.0-rc.15“latest” (prismaCli()); the DB schema is reset per cell (resetDb()) so results are order-independent.
4. Read these instrument caveats before trusting a metric (found this session)
dead(raw ts-prune) over-counts GS ~10× — it flagsindex.tsbarrel / public-API re-exports as dead because consumers import from source, not the barrel. Use the newdead_realcolumn (barrel-excluded) for the GS-vs-naive contrast. On the hosted arm the raw “GS 23 vs naive 1” collapsed to real “GS 0 vs naive 0”. Expect the same on qwen.layerheuristic was widened to also scan entrypoint files (app|server|index|main.ts) with an entity-only pattern, so it can finally catch weak-rung naive monoliths that access domain data straight from the entrypoint (the previous route-file-only filter scored them 0). This matters most for qwen — the whole point of the layer sub-prediction. Startup boilerplate likepool.query('SELECT 1')is deliberately not counted.dup: GS’s higher duplication is per-entity boilerplate symmetry (parallel repos/routes), not logic clones — a genuine but mild GS cost, not an artifact.testcounts test files only — not executed, not asserted, no mutation score. “More tests” ≠ verified quality.- Oracle strictness / where GS actually loses (forensic). The GS oracle deficit on this arm is not broad — it decomposes into exactly two misses. (a) Two cells scored 0 from an over-strict
readingDate(validated as ISO datetime, rejects the spec’s date-only value;DOMAIN_SPEC §3.4saysdate) — GS’s strict-validation discipline overshooting. (b) Every other GS miss isg6_computed_readsonly (§5 budget/occupancy/history);g1–g5pass in every served GS cell, and the failure is identical across all vendors/reps — a deterministic, method-level miss, not model luck (the cascadeuse-cases.mdUC-3/4/5 carry all §5 edges, so it is an implementation divergence, not an information-loss). Expect the same two GS-flavored misses on qwen; watch whether GS’s systematic g6 miss trades against naive’s stochastic g2/rule misses — that contrast is part of the result. Use the median [IQR], not the mean. The exact g6 assertion is now named (re-served gemini-gs/1): P22 —GET /budgeton a paddock with no reading must be422(§5.1), but the GS build returns200 {grazingDaysLeft:null}(conflates “unmeasured” with “measured-but-idle”). 5/6 computed-read assertions pass; only this edge fails, cross-vendor. It is a one-line handler fix — do NOT patch generated cells (would corrupt the sample), just expect this exact g6 signature on qwen-gs too.
5. The final analysis (PREREGISTRATION §7 — do this once qwen is measured)
- Re-run both instruments so
static_cr.json+conformance_cr.jsoncontain all rungs (qwen + the hosted four). Or merge with the hostedcr_analysis_mid_frontier.jsondeltas. - Per (model × condition) cell: median [IQR] of each metric over k.
- Δ(m) per metric, oriented as GS-benefit, across the a-priori ladder
qwen7b → qwen32b → gpt-mid → {frontier trio}. - Test H1: monotone-trend statistic (Jonckheere-Terpstra or Page’s trend test) on Δ across the ordered ladder; report the trend statistic + effect size, not a high-powered p (small k → coarse p; report direction + size, per §7). H1 = Δ declines weak→strong and Δ(frontier) ≪ Δ(weak).
- The money figure: the Δ-vs-capability curve, one line per metric, left→right.
- Report nulls in the same voice as positives (§7). Do not re-order the ladder post hoc (§8); the order in
ladder.jsonis fixed.
6. What the hosted arm already tells you (so you know what to expect)
At mid + frontier the GS oracle benefit is ≤ 0 (naive matches or beats GS) — consistent with the master line “the frontier absorbs the mechanical delta.” If H1 holds, the qwen rungs should show GS’s oracle/conformance advantage turning positive and largest at qwen7b, with Δ declining toward the frontier. If instead Δ is flat/≤0 even at qwen7b, H1 is rejected and the mechanical benefit does not recede — either outcome is a real result (§10). The trust-axis value of GS is deliberately not measured by this experiment.