Thesis Re-Center — Proposal v2 (unified thesis) — SIGNED 2026-09-21

STATUS: SIGNED by JC 2026-09-21, with two amendments now folded in: (i) the stateless reader is central but underlying, not load-bearing (§1); (ii) the trajectory carries no prediction claim — it was reasonable, not foreseen; the speed astonished; the internalization we did not foresee (§5). The Compendium re-center to this arc is authorized.

2026-09-21. Replaces the earlier two-paper split (retired). A proposal to re-center the canonical Compendium and every derivative around ONE unified thesis. Nothing in the Compendium is touched until JC signs off (§10). Grounded in this session: the master line, the discipline-revival model (docs/discipline-revival-model.md), the experiment ledger (docs/white-paper/EXPERIMENT-LEDGER.md), and the NX revival experiment (experiments/nx/).

0. The move — one unified thesis, not two papers

The prior split (“Paper 1 = what receded, Paper 2 = the new bet”) framed the maturation as loss. It is not loss. The original aim was correct programs without writing or reading the implementation — and that aim is met. What grew along the way was the definition of correctness and the discovery of what endures beyond the model. That is one coherent arc, and a stronger, more honest narrative than a list of contributions.

1. The central thesis (the arc)

We set out to produce correct programs that a person ratifies from the specification, the contracts, and the audit trail — without reading the implementation. That aim is achieved. Along the way the meaning of “correct” deepened — from compiles/runs to verified, reproducible, auditable, governable — and the durable core emerged: in a world where the assistant does everything from a single entry point, the value that outlives any model is the guardrail kept OUTSIDE the model and a spec imprint from which the code is regenerable.

  • Underlying premise (central but NOT load-bearing): the stateless reader — the executor derives only from what is externalized. It is the rationale for why each project’s authored internal structure derives the rubric’s attributes; it sits beneath the thesis, it is not the headline contribution. (Framed as derivability, the invariant that survives even as models gain memory; “statelessness” is the vivid case, not the load-bearing term.)
  • Achieved goal: correct programs ratified from the spec/contracts/trace, not the code.
  • What expanded: the definition of correctness. The seven properties are that expanded definition (Executable, Auditable, Defended, … beyond “compiles”).
  • What endures: external guardrails (the model cannot be its own trustworthy verifier) + the regenerable spec imprint (code as residue). This is the differentiator in the single-entry-point world every company is drifting toward.

2. The spine (ordered structure the Compendium adopts)

premise (stateless reader / derivability) → the discipline (GS: the artifacts that make a program derivable and ratifiable) → the expanded definition of correctness (the seven properties) → the mechanism of “how much rigor” (cheap rigor + the revival model — §4) → the durable core (external guardrails + regenerable imprint in the single-entry-point world).

3. The enduring contribution (lead with this)

In the single-entry-point world, two things do not recede:

  1. The guarantee lives outside the model. An LLM cannot be its own trustworthy verifier; the check must be a non-LLM gate. GS is the discipline that makes the artifacts checkable by it.
  2. The intent/decision trace persists outside the session. The model’s reasoning evaporates; the spec, decisions, gates, and ledger persist for governance, and the code regenerates from them. This is the master line, and it is the paper’s durable claim — an argument backed by EX (the gate catching 15 defects against a live system) and the governance framing, presented as a conceptual contribution, not an over-claimed empirical one.

4. Cheap rigor + the revival model — the falsifiable mechanism (kept inside, the scientific risk)

As correctness deepens, the live question becomes how much rigor, which practices, for this project. Cheap rigor answers it economically (the executor revives practices abandoned for cost), and the revival model answers it calculably (a project → the portfolio of practices worth applying, what each buys). This is the paper’s falsifiable spine — the part that can be tested and can fail (the risk JC wants), positioned INSIDE the unified thesis, not as a separate paper.

  • Instances/evidence: Loom (formal spec → compiler, the extreme revival), NX (N-version: revival is exposure-gated not cost-gated; the value is a capability hump, not monotone; and it surfaces generator uncertainty), CR (structural cleanliness, receding). The key finding: a practice’s revival value = f(capability gap), and its SHAPE depends on the practice class — governance = flat/durable, defect-catching = hump, structural = receding.
  • Novelty positioning (avoid the “obvious” trap): not “AI is cheap” but the class characterization (validated-but-dead-on-cost), the predictive model, its validation, and the surprising negatives (which stay dead, and why).

5. Research residue / legitimacy (the honest trajectory)

The early findings the frontier has since absorbed — the bridge (better order), sentinel (retrieval), phase collapse (correctness reconstituted by gates), even better-than-naive code — are reported as a documented trajectory of what we measured, with three honest framings and no claim of prediction (we registered no forecast anywhere — do not say or imply “we called it”):

  1. It was reasonable that the frontier would solve these problems — the levers were real, which is why they were worth naming; that the frontier absorbed them confirms they were real levers.
  2. What was astonishing was the SPEED — how fast the absorption happened.
  3. The direction we did NOT foresee: that the field would move to solve everything internally, from a single entry point, without an external harness or guardrails. We did not predict the internalization. And it is precisely that unforeseen internalization that makes our position durable: when everything moves inside one entry point that is also its own verifier, the guardrail kept OUTSIDE the model and the regenerable spec imprint become the differentiator. This is a section of the arc — the honest record of what we measured and what surprised us — not a headline, and never a prediction claim.

6. Naming and scope discipline

  • Do not coin a new paradigm — that is the “grandiose coinage” reviewers punish; great paradigms are named by the community after adoption. Keep “Generative Specification” as the method’s name; use “cheap rigor” as the phenomenon’s frame-name; anchor the position to the known lineage (declarative / pragmatic-tier / specification-driven).
  • Sharpen “without reading code” everywhere to “ratified from the spec/contracts/trace, without reading the implementation.”
  • One spine, residue relegated (§5). Resist the over-bundling the reviewers flagged: the paper has one arc, not five competing contributions.

7. What changes in each derivative

  • Compendium (canonical master): re-order to the §2 spine. Keep §III (stateless reader/derivability) as the premise. Re-frame the seven properties as the expanded definition of correctness. Add cheap rigor + the revival model as the mechanism (§4), with Loom/NX/CR. Elevate the durable core (§3) to the lead durable claim. Fold the early findings into a trajectory section (§5).
  • IEEE: this becomes the paper — re-centered to the arc; the contribution list = derivability + expanded-correctness/governance (durable) + the revival model (falsifiable mechanism), with the trajectory as the honest longitudinal record. Related work now includes the revived practices’ literature (formal methods, N-version/Avizienis, Cleanroom, PBR/Basili) — filling the 13→30-50 gap.
  • WhitePaper / FieldGuide: lead with the durable core + “how much rigor does YOUR project need” (the revival model = the product’s Assessment output; paper and product tell one story).
  • Course (GS Core): a canon change → a G-row to the course board; C15 (“what you’re buying”) hosts the durable-core + cheap-rigor reframe.

8. Experiments — evidenced now vs owed

  • Evidenced today: derivability (RX), retrieval (KX), governance/gate (EX), the seven properties (BX), the early wins + saturation (AX), the receding (AX2/CR), and the first revival datum (NX).
  • Owed (the risk): generalize revival to a 2nd/3rd practice (mutation, PBR, formal-spec) with harder problems (λ>0 on the frontier); calibrate/validate the model against measured outcomes (the chronicle-ledger dataset); human validation (DX2 redesigned). The paper states the revival model as model + hypothesis with instances, promoted to a proven law only as these land.

9. Risks and guardrails (honest)

  • Over-bundling (reviewers’ own flag) → §6 scope discipline.
  • Legitimacy-from-absorption unfalsifiable if asserted → §5 anchor to dated experiments only.
  • “Without reading code” over-claim → §6 scoping.
  • Cheap-rigor grandiosity → §4 novelty positioning + the owed experiments before “law”.

10. Decision requested

Sign off on: (a) the unified thesis and the §2 spine; (b) the durable core as the lead contribution (§3); (c) cheap rigor + the revival model as the falsifiable mechanism inside (§4); (d) the early findings as documented trajectory/legitimacy anchored to dated experiments (§5); (e) the naming/scope discipline (§6). On sign-off, I rewrite the Compendium §-structure to this arc and propagate to the derivatives in one deliberate pass.