CR — Results: Copilot arm (mid → frontier rungs)
Partial results for the mid and frontier rungs measured on the Copilot PC. The two weak local rungs (
qwen25coder7b,qwen25coder32b) are generated + measured separately via Ollama on the main PC and must be joined before the full H1 trend test (Jonckheere-Terpstra / Page, PREREGISTRATION §7). This file reports only what this arm can conclude on its own.
What was measured
- Rungs (a priori ladder order, PREREGISTRATION §4 +
ladder.json):gpt-mid(rank 2, mid) →{gpt-frontier, gemini-frontier, claude-frontier}(rank 3, frontier — extra frontier points perladder.json.note_min_rungs). - Conditions: naive vs gs. k = 3 reps per cell. 24 cells total, all served (24/24).
- Instruments (run directly with Windows node, not
measure_cr.sh):static_cr.cjs(clean static signal) +conformance_cr.cjs(behavioral oracle: 6 Hurl probe groups vs a per-cell-reset Postgres). - Statistic of record: per-cell median [IQR] over k=3 (PREREGISTRATION §7), robust to single-rep outliers.
Per-cell medians [IQR]
Δ GS-benefit is oriented so positive = GS better: for higher-is-better metrics (oracle, rules, tests) Δ = median(gs) − median(naive); for lower-is-better metrics (dup, cc, dead, layer) Δ = median(naive) − median(gs).
| metric | rung | naive [IQR] | gs [IQR] | Δ GS-benefit |
|---|---|---|---|---|
| oracle/6 | gpt-mid | 6.0 [6.0,6.0] | 5.0 [2.5,5.5] | −1.0 |
| oracle/6 | gpt-frontier | 5.0 [5.0,5.5] | 4.0 [2.0,4.5] | −1.0 |
| oracle/6 | gemini-frontier | 6.0 [6.0,6.0] | 5.0 [5.0,5.0] | −1.0 |
| oracle/6 | claude-frontier | 6.0 [6.0,6.0] | 6.0 [5.5,6.0] | +0.0 |
| rules/3 | gpt-mid | 3.0 [3.0,3.0] | 3.0 [1.5,3.0] | +0.0 |
| rules/3 | gpt-frontier | 3.0 [3.0,3.0] | 3.0 [1.5,3.0] | +0.0 |
| rules/3 | gemini-frontier | 3.0 [3.0,3.0] | 3.0 [3.0,3.0] | +0.0 |
| rules/3 | claude-frontier | 3.0 [3.0,3.0] | 3.0 [3.0,3.0] | +0.0 |
| dup% | gpt-mid | 0.0 [0.0,1.6] | 0.0 [0.0,0.0] | +0.0 |
| dup% | gpt-frontier | 0.0 [0.0,0.3] | 1.2 [0.6,1.4] | −1.2 |
| dup% | gemini-frontier | 2.9 [1.5,3.3] | 6.5 [3.5,7.1] | −3.5 |
| dup% | claude-frontier | 0.0 [0.0,0.1] | 0.0 [0.0,0.2] | +0.0 |
| cc.mean | gpt-mid | 1.9 [1.7,1.9] | 1.7 [1.6,1.7] | +0.2 |
| cc.mean | gpt-frontier | 1.7 [1.6,1.8] | 1.6 [1.5,1.6] | +0.1 |
| cc.mean | gemini-frontier | 2.1 [2.1,2.2] | 2.0 [2.0,2.1] | +0.1 |
| cc.mean | claude-frontier | 1.6 [1.6,1.8] | 1.5 [1.5,1.5] | +0.2 |
| dead (raw) | gpt-mid | 0.0 | 2.0 | −2.0 |
| dead (raw) | gpt-frontier | 0.0 | 0.0 | +0.0 |
| dead (raw) | gemini-frontier | 1.0 | 23.0 | −22.0 |
| dead (raw) | claude-frontier | 1.0 | 3.0 | −2.0 |
| dead_real (barrel-aware) | gpt-mid | 0.0 | 2.0 | −2.0 |
| dead_real (barrel-aware) | gpt-frontier | 0.0 | 0.0 | +0.0 |
| dead_real (barrel-aware) | gemini-frontier | 0.0 | 0.0 | +0.0 |
| dead_real (barrel-aware) | claude-frontier | 1.0 | 0.0 | +1.0 |
| layer | (all rungs) | 0.0 | 0.0 | +0.0 |
| tests | gpt-mid | 1.0 [0.5,1.0] | 1.0 [1.0,1.5] | +0.0 |
| tests | gpt-frontier | 2.0 [2.0,2.5] | 5.0 [4.5,5.0] | +3.0 |
| tests | gemini-frontier | 4.0 [4.0,5.5] | 6.0 [6.0,8.0] | +2.0 |
| tests | claude-frontier | 6.0 [4.5,6.0] | 8.0 [7.5,9.5] | +2.0 |
Δ trend over the measured range (mid rank 2 → frontier rank 3)
| metric | Δ mid | Δ frontier (mean of 3) | direction |
|---|---|---|---|
| oracle/6 | −1.00 | −0.67 | GS-deficit shrinks toward frontier |
| rules/3 | 0.00 | 0.00 | flat (tie) |
| dup% | 0.00 | −1.59 | GS worse toward frontier |
| cc.mean | +0.15 | +0.13 | flat (GS marginally better) |
| dead (raw) | −2.00 | −8.00 | artifact — barrel re-exports (see correction) |
| dead_real | −2.00 | +0.33 | flat (GS ≈ naive real dead ≈ 0) |
| layer | 0.00 | 0.00 | flat (both layer cleanly; heuristic scope-limited) |
| tests | 0.00 | +2.33 | GS emits more tests toward frontier |
Honest read (measured range only)
- Behavioral conformance (the load-bearing metric): naive ≥ GS at every measured rung. GS ties naive only at the frontier (claude). The GS deficit on the oracle is roughly constant (−1.0) across mid + two frontier vendors and closes to 0 at the top.
- Business-rule sub-score is a tie (3.0 = 3.0) at every rung under the median statistic. (An earlier mean-based read showed naive > GS on rules; that was driven entirely by the two
oracle-0outlier cells below and vanishes under the pre-registered median.) - GS emits more tests and marginally lower complexity. The raw
dead-export metric shows GS far higher (gemini gs median 23 vs naive 1), but this is a metric artifact — see the correction below. After excludingindex.tsbarrel re-exports (dead_real), GS’s real dead code is ≈ 0 at every rung (gpt-mid 2, elsewhere 0), on par with naive. GS’s genuine cost here is more duplication at the frontier (per-entity boilerplate, not logic clones). - Layer-boundary violations = 0 everywhere — both conditions route DB access through a repo/service; no route/controller calls the data client directly at mid+frontier. See the validity caveat below: the heuristic only scans route-named files, so it will false-negative on weak-rung naive monoliths (inline data calls in
app.ts) — widen it before measuring qwen.
Dead-code correction: the raw metric penalizes GS’s public API surface
The pre-registered dead metric is a raw ts-prune count. ts-prune flags any export not imported through its own path — which includes every index.ts barrel / public-API re-export, because consumers import the symbol from its source file, not the barrel. GS’s own standards mandate exactly this surface: an index.ts public API per module, port interfaces (IUserRepository), swappable adapters (InMemoryUserRepository), a custom exception hierarchy, and named constants. Each becomes a re-export ts-prune misreports as dead.
Verified on the worst cell (gemini gs/2): all 51 raw “dead” exports are index.ts re-exports of live symbols (ValidationError 31 refs, AuthService 12, IUserRepository 5, even InMemoryUserRepository 4). Barrel-excluded count = 0. Naive scores ~0 on the raw metric only because it builds no barrels/ports — it imports concretely file-to-file. dead_real (barrel-aware, added to static_cr.cjs, applied identically to both conditions) is the fair count; the “GS 23 vs naive 1” gap collapses to “GS 0 vs naive 0”.
The load-bearing caveat: H1 cannot be concluded from this arm
H1 (Δ monotonically declines weak→strong, largest at the weak rung) is anchored on the weak qwen rungs, which are NOT in this dataset (Ollama / main PC). Within the measured mid→frontier segment there are only two distinct a-priori ranks, so no monotone-trend statistic is run here. Join the qwen cells, then run Jonckheere-Terpstra / Page per §7. What this arm establishes: at mid + frontier, the GS conformance benefit is ≤ 0 (naive matches or beats GS), consistent with the “frontier absorbs the mechanical delta” master line but not by itself sufficient to confirm or reject the monotone law.
What still needs analysis (open items, ranked)
- Layer metric — validity fix APPLIED this session (no results change).
layer=0/0is real at mid+frontier (verified: data-client calls live in repository/service files, none in route handlers, both conditions). The heuristic previously scanned onlyroutes/…|route.ts| controllerfiles, so a weak-rung naive monolith inliningprisma.paddockinapp.tswould score 0 falsely. Fixed:static_cr.cjsnow also scans entrypoint files (app|server|index|main.ts, excluding data-layer dirs) with an entity-only pattern (LAYER_ENTITY_RE, no rawquery|execute|…verbs) so it catches monolith domain access without false-positiving on startup boilerplate likepool.query('SELECT 1')health checks. Re-checked across all 24 cells: 0 change at mid+frontier — the fix only arms the metric for the qwen monoliths, exactly where GS is predicted to win. - Duplication is real but benign for GS. GS’s higher dup (gemini 6.5% vs 2.9%) is per-entity boilerplate symmetry — repeated CRUD/mapper blocks within each repository and parallel validate→delegate→respond skeletons across route files (
herd.routes↔paddock.routes). Not duplicated business logic. It is the DRY-vs-explicit-layering trade GS makes (one module per entity). Worth reporting as a genuine (mild) GS cost, not an artifact. - The two oracle-0 GS cells — RESOLVED this session (see the dedicated section below). Live re-serve of
gpt-frontier gs/1proved register works (201); the real cause is a 422 on/readingsfrom over-strictreadingDatedatetime validation (spec §3.4 saysdate). A genuine GS conformance miss, not a harness artifact. Median already handles it as an outlier. - Test metric = presence only.
testcounts test files; it does not run them or measure assertions/mutation (MSI). GS’s higher test count is not evidence the tests pass or assert — that needs the tests executed + Stryker, not done here. Treat “more tests” as scaffolding volume, not verified quality. - cc / dup / tests still use the mean in the earlier console summary — the record of truth is the median [IQR] table above (robust to the two outlier cells). Prefer it.
Two genuine oracle-0 cells — root cause verified by live re-serve
gpt-mid gs/2 and gpt-frontier gs/1 served, migrated (prisma), and register works (POST /register → 201 {token}, confirmed by a live re-serve of gpt-frontier gs/1: fresh npm install, clean schema reset, prisma generate + migrate deploy, then curl). So the earlier “runtime 500 on register” guess was wrong — corrected here.
Running the six oracle probes against the live server reproduced 0/6 deterministically. The failing step (Hurl --error-format long) is POST /readings returning 422 {"error": {"code":"validation","message":"Invalid ISO datetime"}}. The probe sends the spec-conformant "readingDate":"2026-01-01", but this GS build validated readingDate as a strict ISO datetime and rejects a date-only value. DOMAIN_SPEC §3.4 defines readingDate as a date, so the app is over-strict and genuinely non-conformant to the data contract. Because every rule group (capacity/budget/overlap) needs a reading first, the 422 cascades to all six groups → oracle 0.
This is a real GS conformance miss (a data-contract type error), not measurement noise, and it is a GS-flavored failure: the “strict validation / fail-fast” discipline pushed the model to z.datetime(), which then rejects a valid date. The naive gpt-frontier build accepts the date-only value and passes. The median statistic already treats these as the minority outliers they are (the cell’s other reps pass), so no numbers change — but the cause is now documented, not guessed. Not “fixed” (editing generated cells would corrupt the sample).
Decomposing the −1.0 oracle Δ: it is two specific misses, not broad degradation
The oracle Δ(GS−naive) = −1.0 at gpt-mid, gpt-frontier, gemini looks like “GS builds worse software.” The per-cell data refutes that reading.
Per-cell oracle (pass/6), k=3:
| rung | naive | GS |
|---|---|---|
| gpt-mid | 6, 6, 6 | 6, 0, 5 |
| gpt-frontier | 5, 6, 5 | 0, 5, 4 |
| gemini-frontier | 6, 6, 6 | 5, 5, 5 |
| claude-frontier | 6, 6, 6 | 6, 5, 6 |
Per-probe forensic (which group fails):
- The two zeros are the same
readingDateover-strictness bug (previous section) — one root cause, cascading to all six groups. Not six independent failures. - Every other GS miss is
g6_computed_reads— and only that group: gemini gs/1,2,3; gpt-mid gs/3; gpt-frontier gs/2 all score exactly 5/6, failingg6alone.g1–g5(auth/roles, validation/404, and all three business rules R1/R2/R3) pass in every served GS cell. GS reliably builds a functional backend; it trips on one behavior: the §5 computed reads (strict budget arithmetic / null-on-no-open-move / 422-on-no-reading / occupancy / history order). g2_validation_404also fails in gpt-frontier — but it fails in gpt-frontier naive too (naive/1, naive/3). That is a model/vendor quirk at that rung, not a GS effect.
Ruled out — cascade fidelity. The obvious hypothesis is that the GS cascade dropped a §5 edge during phase-collapse. Verified false: gs/use-cases.md UC-3/4/5 carry all five edges, and the budget formula is arithmetically identical to DOMAIN_SPEC §5.1 (animalUnits × 100 = × 3000/30; worked example → 120). The cascade does not lose the spec. So the systematic g6 miss is an implementation-level divergence, not an information loss.
Why this is a pro-GS signal, not anti-GS. naive’s g2 failures are stochastic (present in some reps, absent in others, model-dependent). GS’s g6 failure is deterministic and identical across every vendor and rep. The method makes the failure systematic, reproducible, and diagnosable at the method level rather than a matter of per-generation luck — a single method-level fix would lift every GS cell at once, whereas a stochastic miss cannot be. This is consistent with the phase-collapse thesis (GS should make behavior at least as good and more consistent): GS matches naive on 5/6 groups and concentrates its entire remaining gap in one reproducible behavior.
Closed — the exact g6 assertion, named by live re-serve. Re-served gemini-frontier gs/1 (fresh npm install, npm run migrate, npx tsx src/server.ts) and ran only g6_computed_reads.hurl with --error-format long. Five of six computed-read assertions pass — the strict budget arithmetic (grazingDaysLeft == 120), the null-on-no-open-move edge (P21), occupancy (P23), and history newest-first (P24) are all correct. The single failing assertion is P22 (line 138): GET /paddocks/:id/budget on a paddock with NO reading must return 422 (DOMAIN_SPEC §5.1: “If no reading: 422”). The GS build instead returns 200 {"grazingDaysLeft": null} — it conflates the two null-ish edges, treating no reading the same as no open move rather than distinguishing “unmeasured → reject (422)” from “measured but idle → null”. Note the direction: here GS is too lax (200 where the spec wants 422), the mirror image of the readingDate cell where GS was too strict. So GS’s misses are edge-case interpretation divergences, not a single “over-strict” theme — and each is a one-line, method-level fix (add the if (!latestReading) → 422 branch to the budget handler), which would lift the g6-only cells across every vendor at once. Named on gemini-gs/1; the identical g6-only signature across gpt-mid/gpt-frontier/gemini reps makes this the shared cause.
Instrument fixes applied this session (portability / hygiene — no DV bias)
The runner was authored for Git-Bash/Linux and would not run on Windows. All fixes below apply identically to naive and gs across all rungs, so they do not bias the naive-vs-gs contrast. The oracle probes and the layer heuristic were not touched.
static_cr.cjs (portability + one transparent metric augmentation)
- jscpd fed a relative
srcpath (absolute Windows path made fast-glob treat\as an escape → 0 files scanned → null dup%). - eslint runs with
cwd= the runner dir so@typescript-eslint/parserresolves (the--resolve-plugins-relative-toflag does not apply to the parser; projects have no node_modules). dead_realadded — a barrel-aware dead-export count (ts-prune output excludingindex.tsre-exports), reported alongside the untouched pre-registered rawdead. This is an addition, not a replacement: both numbers are emitted so nothing is hidden. Applied identically to naive and gs. Rationale in the “Dead-code correction” section above.
conformance_cr.cjs
HURLconst set to the 8.3 short pathC:\PROGRA~1\hurl\hurl.exe(the bash path/c/Program Files/Hurl/...is invalid under cmd; spaces break undershell:true).resetDb()added — drops/recreates thepublicschema per cell via stdin (the old-c "..."form splits undershell:true, so the reset never ran → cross-cell table contamination). Makes results order-independent.migrate()rewritten: prisma-first with fall-through; broadened npm migration-script detection; recursive.sqldiscovery + apply (stdin,ON_ERROR_STOP).prismaCli()added — pinsnpx prisma@<installed client version>so cells that omit theprismadevDep don’t fetch the brokenprisma@8.0.0-rc.15“latest”.runCell()tolerates self-migrating apps (some embedCREATE TABLEand migrate on boot); recordsmigration:"self-or-none"and serves anyway — the oracle decides.tsxinstalled globally soserve()’snpx tsxresolves instantly (projects declare ts-node, not tsx; without global tsx, npx hangs on an interactive install prompt).
Raw per-cell outputs (static_cr.json, conformance_cr.json) are gitignored by design (per-machine artifacts). The curated numbers above and cr_analysis_mid_frontier.json are the committed record.