IV. GENERATIVE SPECIFICATION: THE DISCIPLINE

Draft for IEEE Access. Parent section IV. Introduces the seven properties and the mechanisms. The subsection IV.A, “Structure of the Seven” (companion file 02-section-IVa-structure-of-the-seven.md) analyzes the internal structure of the seven and is not reproduced here; it follows this section’s lead enumeration. The mechanism subsections below are lettered IV.B through IV.E to leave IV.A to that companion.

House constraints honored: seven-property instrument only; token economics qualitative; objective metrics are the evidence of record; no em-dashes.

Sections I through III established why a new discipline is necessary and stated the binding constraint precisely: the AI executor is a stateless reader that begins each session with no memory of prior sessions, no institutional context, and no ability to ask clarifying questions, and everything not present in the artifacts is absent. The binding constraint that follows is derivability: a system’s lifecycle layer is derivable when a stateless reader, given its artifact set alone, can correctly determine what should be built, where, why, and to what contracts, without external human context. This section defines the discipline that satisfies that constraint. It consists of a small set of specification properties that operationalize derivability, a lexical distinction that removes the degrees of freedom a stateless reader would otherwise fill arbitrarily, and two structural consequences that follow once the specification is complete and the executor is capable.

Following R. C. Martin’s sense of a paradigm as a discipline defined by what it removes from programmer freedom, Generative Specification removes the freedom to leave architectural intent implicit [Martin, 2002]. The removal is absolute for its reader, because the stateless executor has none of the recovery mechanisms (memory, collaboration, accumulated convention) through which a human reader compensates for an underspecified requirement. What prior paradigms made inconvenient, this discipline makes structurally absent.

The seven properties below are to Generative Specification what SOLID is to object-oriented programming: a named, teachable set of obligations that makes the discipline concrete, inspectable, and transferable. Each property names a specific failure mode observed in production, and together they operationalize derivability by making each class of failure structurally unreachable. Table I states each property, its one-line definition, and the failure mode it removes.

TABLE I. The Seven Specification Properties.

Property Definition Failure mode removed
Self-describing The artifact set explains its own architecture, decisions, and conventions from its own contents; no external knowledge is required. Session amnesia: the reader cannot recover what the system is or why, and completes the gap arbitrarily.
Bounded Every unit of work declares an explicit scope and seams, and stays within the reader’s finite read budget. Scope leakage and silent truncation: the reader edits against an incomplete or unbounded view.
Verifiable The correctness of any output can be checked without human judgment. Divergent definitions of done: write-completion mistaken for semantic validity.
Defended Destructive or forbidden operations are structurally prevented, not merely discouraged. Aspirational controls that never run, so a known-forbidden pattern still reaches production.
Auditable The current state and the history that produced it are fully recoverable from the artifacts alone. Intentional tradeoffs read as defects and “corrected” into silent drift.
Composable Units combine and extend without unexpected coupling, and each is navigable in isolation. Propagation surprise: a local change breaks distant callers the reader cannot locate.
Executable The generated output satisfies its behavioral contracts against a real execution environment, not merely compiles and passes static analysis. Fully static-valid output that fails every integration test against a live system.

Two of the seven, Self-describing and Bounded, carry disproportionate weight, and the reason is mechanical rather than stylistic. A large language model does not lack the knowledge to write correct code; its training corpus already contains the formal tradition a correct implementation draws on, from Hoare logic [Hoare, 1969] and type theory to design by contract [Meyer, 1992] and the SOLID principles. What the model lacks at the moment of generation is the instruction of which region of that knowledge to apply here. This has a precise analog in cognitive science. Bransford and Johnson (1972) showed that a context label supplied before an ambiguous passage (“laundry,” “music”) made it immediately comprehensible, because the label activated the reader’s pre-existing schema and supplied the framework that filled every ambiguous phrase with correct meaning. The same mechanism operates in generative development with the direction inverted: where schema fit aids comprehension, a specification cue aids generation. A specification that is Bounded (each unit declares a narrow scope and loads only its own slice of the system) and Self-describing (each unit announces what it is and what it is for) selects the correct subset of the model’s trained knowledge and suppresses adjacent patterns from its prior distribution. A specification that fails to bound its scope or describe its identity activates the wrong schema, or none, and the downstream derivation is arbitrary regardless of how carefully the remaining five properties are satisfied. This is why constraint, counter-intuitively, does not weaken the generator: it selects the path the generator was already capable of taking and prunes the ones the specification declares wrong. Subsection IV.A analyzes the internal structure of the seven, including the partition of the remaining five properties, and is not repeated here.

The seven properties were operationalized as a fourteen-point rubric (0, 1, or 2 per property) that served as the internal measurement instrument during development of the discipline. The evaluation of record in this paper, however, is reported through objective, rubric-independent metrics (mutation score, real coverage, and static-analysis counts) and a blind adversarial audit, presented in Sections V and VI.

IV.B. Prescriptive Versus Descriptive Specification

The seven properties state what a derivable specification must satisfy; a single lexical distinction determines whether a given clause satisfies them. A specification can be written descriptively or prescriptively, and the difference is decisive for a stateless reader. A descriptive obligation records a state of affairs and leaves the reader to resolve the ambiguity it contains. Given “the actor needs to see certain data,” the stateless reader resolves toward the literal minimum, because nothing in the artifact rules the minimum out, and every unstated assumption is a degree of freedom the reader fills arbitrarily, at generation speed. A prescriptive obligation removes that degree of freedom. Given “the dashboard MUST return counts aggregated by type, MUST compute active members in the last seven days, and MAY cache the result,” the reader has no ambiguity left to resolve.

The discipline phrases obligations with the standard normative vocabulary of RFC 2119 as clarified by RFC 8174 [Bradner, 1997; Leiba, 2017]: MUST, MUST NOT, REQUIRED, SHALL for hard obligation, SHOULD, SHOULD NOT, RECOMMENDED for defeasible obligation where deviation requires a recorded reason, and MAY, OPTIONAL for the permitted. Per RFC 8174 the words carry normative force only when capitalized, leaving ordinary prose unaffected. This closed vocabulary serves Bounded and Self-describing directly, because obligation level is then parsed deterministically rather than inferred, and it binds each clause to two of the other properties at once. Every MUST is an acceptance criterion, and therefore a verification probe (Verifiable); the keyword also sets the severity of the gate that enforces it, so that MUST maps to a blocking gate, SHOULD to a warning, and MAY to an ungated permission (Defended). The discipline is selective rather than total: the load-bearing obligations (acceptance criteria, constraints, and prohibitions) are keyworded, not every sentence, because over-marking is the same excess that degrades any bounded artifact.

IV.C. Phase Collapse

A further mechanism the discipline enables is phase collapse, a named phenomenon in which the classically separate activities of a development cycle converge into a single derivation step. In conventional practice, planning, implementation, verification, review, and deployment are distinct phases separated in time, each with its own artifacts and handoffs. Under Generative Specification, when the specification is complete and the executor is a capable AI reader, these phases collapse into one derivation: the specification holds the full intent, the executor derives the implementation and its tests together, and the quality gates close the loop before the session ends. The phases do not disappear as obligations; they cease to be temporally separate moments.

The collapse is a direct consequence of the reader. A human workflow enforced phase separation through the passage of time and the movement of work between people, and several disciplines depended on that separation. Test-driven development, for instance, requires a strict sequence of failing test, confirmed failure, then implementation. A generative agent in a single context window has no temporal barrier between test authorship and implementation authorship, so the RED phase ceases to exist as a distinct moment. This RED-phase collapse is the function-level instance of the general phenomenon [after Collins, Brown, and Newman, 1989, on making tacit sequence explicit]. Where phase separation formerly did the enforcing, structural gates must now do it: a specification whose phases have collapsed remains correct only because the Verifiable and Defended properties reconstitute, as gates, the guarantees that temporal separation once provided.

IV.D. Cost Inversion and Tokens per Correct Output

Phase collapse changes the economics of iteration. In a traditional cycle, wrong output costs a sprint: read the code, locate the error, write a correction, review, wait for CI, and merge, with coordination overhead accumulating at every boundary. Under this discipline, wrong output costs one sentence: identify the missing constraint, add it to the specification, and regenerate, because the stateless reader re-reads the complete specification and derives from scratch. As the cost of regeneration approaches zero, the relationship between code and specification inverts. The code becomes residue, a regenerable output of the specification rather than the durable asset, and the specification becomes the scarce good, because it is the only artifact that carries intent across the session boundary. The residual cost is real but bounded by the practitioner’s own domain fluency, namely the cost of identifying which constraint is absent, not by organizational friction.

This inversion also reframes the most common objection to specification-first development, that authoring a specification and a navigation tree spends tokens ad-hoc prompting does not. The objection measures the wrong quantity. The binding metric is not tokens generated but tokens per correct, accepted output. A session without authored structure re-reads the codebase and re-derives its architecture on every invocation, so the stateless reader pays the discovery cost repeatedly; a session with a bounded, self-describing specification reads the authored structure once and navigates it directly. Independent measurement outside this domain reports the same principle in retrieval: reading a pre-authored compact knowledge structure costs substantially fewer tokens at higher accuracy than re-deriving structure from prose at query time [Yarmoluk and McCreary, 2026]. We therefore make the economic claim qualitatively and hold the direction, not any specific magnitude, as the load-bearing point: when regeneration is near-free, the specification is where value concentrates, and the correct efficiency metric is tokens per correct output.

IV.E. The Loop: Retrieve, Generate, Verify

The discipline runs as a three-step loop, and naming its steps separates the two concerns that the token objection above conflates.

Retrieve. The loop begins by assembling the correct context for the stateless reader. Without authored structure this is a retrieval problem in the literal sense: the reader scans, embeds, or re-derives the architecture from the code on every session. Generative Specification attacks the retrieval from both ends. It authors the structure, so that a sentinel navigational tree (a lossless hierarchy of specification files, each declaring its own scope and routing to its children, with the root always loaded within the bounded line budget) is traversed by direct lookup rather than similarity search, mitigating both the token cost and the mid-context accuracy degradation reported for long windows [Liu et al., 2023]. It also shapes the code to be retrievable, because the structural disciplines make predictable placement, typed contracts, and intentional naming the conditions under which the reader’s first search succeeds. This step realizes the Self-describing, Bounded, Composable, and Auditable properties.

Generate. The second step is the derivation that prompting optimizes, and it is the one phase-collapse compresses. Given a complete grammar, derivation is mechanical, not because the model is intelligent but because completeness has closed the space of valid outputs to those that are correct.

Verify. The third step is generative execution: checking generated output against the specification in a real execution environment, not assembling context. This is categorically distinct from retrieval. It runs the full test pyramid (unit, integration, and end-to-end tests, extended with mutation testing where applicable) and, above it, AI-as-QA confirmation, in which a vision-capable executor exercises the running application against the specification’s acceptance criteria and returns a structured gap analysis. This step is what realizes the Executable property, which the mere existence of checks (the Verifiable property) does not: a system can be fully Verifiable, with correct types, passing lint, and well-structured tests, while producing a server that fails every integration test against a live database. The loop closes when the gap analysis is empty. When it is not, the output is not patched in place; the absent constraint is identified, written into the specification, and the loop is re-run, which is the cost-inversion dynamic of Subsection IV.D operating one iteration at a time. In this framing Generative Specification is retrieval-augmented and verified generation, with the retrieval authored rather than inferred.