III. PROBLEM FORMALIZATION
Reviewer-safe framing. Seven-property paper only. No decagon, no SAVED, no CORE, no maturity levels, no 0-100 tool, no token or cost percentages, no confidential field data. House voice: no em-dashes, no semicolons, precise academic English. Citations are [Author, year] placeholders to be renumbered in the IEEE two-column assembly.
III. PROBLEM FORMALIZATION
The dominant failure mode of AI-assisted development is architectural drift: the multi-session accumulation of output that is locally valid and passes its immediate tests, yet is architecturally incoherent with the system it joins. This section states the structural cause of that failure and derives from it the constraint the discipline must satisfy. The argument proceeds in five steps. First we characterize the reader that AI-assisted development now targets, and show that its defining property is a binding design constraint rather than an incidental limitation. Second we state the obligation that constraint imposes, which we term derivability. Third we locate that obligation in the semiotic tripartition of [Morris, 1938] and show that it occupies a tier prior programming disciplines left vacant. Fourth we give the mechanism that explains why externalizing intent for such a reader succeeds, which we call the bridge, and identify its asymmetry as the source of leverage. Fifth we address the failure mode the constraint introduces at scale, context degradation, and state the structural response, the sentinel navigational tree.
A. The Stateless Reader as a Binding Constraint
We define the stateless reader as an executor that begins each session with no memory of prior sessions, no institutional context, no accumulated conventions, and no channel through which to ask a clarifying question. Everything not present in the artifacts is, for this reader, absent. A large language model deployed as a code-generating agent is such a reader. Its context window resets at the session boundary, it carries no persistent memory of decisions taken in earlier sessions [Tulving, 1972], [Squire, 1987], and it has no mechanism for requesting the interpretive context a human collaborator would supply on request.
This property separates the AI executor from the human engineer along a single decisive axis. Where a human engineer interprets an underspecified requirement, compensating across the gap with memory, inference, and accumulated shared context, the AI executor processes what is present. The human compensation layer does not exist. An incomplete description is therefore not deferred until it can be clarified. It is completed immediately, at generation speed, and consistently in whatever direction the model’s prior distribution favors, which need not be the direction the author intended.
The consequence is drift, the software-generation analog of architectural erosion [De Silva and Balasubramaniam, 2012]. Because each session inherits the artifacts produced by the last, an incoherence introduced once propagates to every subsequent session that reads the corrupted context. This is Lehman’s observation that complexity increases unless active work is done to reduce it [Lehman, 1980], operating in the specific case where no specification constraint is present to arrest it. Independent measurement corroborates the mechanism: [Orlanski et al., 2026] report that structural erosion rises across the large majority of AI-agent trajectories and that prompt-level intervention improves initial quality without halting the degradation.
The constraint is not relaxed by larger context windows. An unbounded window over an underspecified codebase is not unbounded derivability. It is an unbounded drift surface. The model reads more of the implicit record, but it cannot derive intent that was never externalized in the first place. The stateless reader is therefore a binding design constraint, not a transient limitation of current model capacity, and the remainder of this section treats it as such.
B. Derivability as the Obligation
The constraint of Section III-A yields a single obligation. We term it derivability. A system’s lifecycle layer is derivable when a stateless reader, given its artifact set alone, can correctly determine what should be built, where, why, and to what contracts, without requiring external human context. A specification satisfies the obligation when it makes the correct output derivable from the artifacts alone.
Correct here carries a broader sense than syntactic well-formedness. A derived implementation state is admissible when it is both structurally well-formed under the specification’s rules and conformant to its behavioral and acceptance-test obligations. A specification whose rules are internally consistent but do not capture the system’s actual obligations is therefore possible, and it is the discipline’s primary failure mode: a specification faces the same verification burden as the implementation it governs.
The obligation has clear antecedents in classical software engineering, which this work instantiates for the AI generation context rather than introduces. [Parnas, 1972] established that a well-decomposed system should make every design decision locatable by inspection, so that a reader with access to the specification can derive the intended behavior without consulting the implementation. [Jackson, 2001] extended this through the Problem Frames approach, requiring the specification to bound the problem sufficiently or the implementation will fill the remaining gap arbitrarily. Derivability is the composition of these two requirements under one reader that has neither memory nor the option to ask. What was, for a human team, a recoverable cost paid by skilled interpretation becomes, for the stateless reader, a structural failure produced at generation speed.
C. Locating the Constraint: The Pragmatic Tier
The derivability obligation occupies a definite place in the semiotic tripartition of signs due to [Morris, 1938], which distinguishes syntactics, the relation of signs to one another, semantics, the relation of signs to what they denote, and pragmatics, the relation of signs to their interpreters in a context of use. Applied to programming discipline, the tripartition sorts existing disciplines by the reader they serve. We use syntactic and semantic in their programming-language senses throughout, which are consistent with the Morris tripartition but distinct from the technical senses these terms carry in formal linguistics.
Syntactic disciplines constrain the form of source artifacts. Structured programming, object-oriented programming, and functional programming, together with structural schemes such as layered clean architecture, govern which constructs are permitted and which dependency directions are allowed. Semantic disciplines constrain the meaning that structure communicates to a human reader who brings interpretive context. SOLID, test-driven development, domain-driven design, and conventional commits all assume a reader with state: colleagues, institutional memory, and shared history. A codebase that violates a semantic discipline still compiles, and its cost is paid by the engineers who must later read it.
Generative Specification operates at the pragmatic tier. It constrains neither the form of what is constructed nor the meaning communicated to a reader who carries context, but what is derivable by a reader who carries no context at all. Every intent that a human team would resolve through shared knowledge must instead be externalized as a formal artifact, because the channel through which shared knowledge travels does not exist for the stateless reader. This tier had no prior occupant, and the reason is specific rather than accidental. Stateless machine readers predate large language models: interface definition languages such as OpenAPI read interface contracts, and formal specification languages such as TLA+ or Alloy verify property invariants. Neither reads the lifecycle layer. An IDL consumer can call an endpoint but cannot determine whether that endpoint should exist. A formal verifier establishes that a design satisfies a stated invariant but does not generate the naming convention, the module boundary, or the decision record an executor must read before implementing, and it cannot reason about which directions of change are valid, because those questions are not expressible in the language it reads. The transformer architecture [Vaswani et al., 2017] produced the first widely deployed reader that reads lifecycle intent expressed in natural language and derives implementation from it. With its deployment, leaving the lifecycle layer implicit changed in kind, from a recoverable cost paid by skilled humans to a derivability collapse.
This last distinction is what separates the pragmatic tier from the semantic tier, and it is one of kind rather than of degree. A semantic violation produces a worse system that a human reader can still recover. A pragmatic violation produces a grammar that a context-free executor cannot parse. The failure is not a quality deficit at higher intensity. It is the absence of a derivation.
D. The Bridge and Its Asymmetry
Sections III-A through III-C establish that intent must be externalized and where the resulting obligation sits. They do not explain why externalizing it succeeds. This subsection advances the mechanism — the bridge and the read-asymmetry — as a conjectured explanation rather than a measured result.
Every structural discipline that engineering already practices, intentional naming, the specification itself, the SOLID interface boundaries, the ubiquitous language of domain-driven design, type-driven design, and design by contract, is a bridge between human conceptual language and executable code. Each encodes human meaning in a form a machine can act on, and machine behavior in a form a human can verify. These disciplines were built to carry intent across that gap for the next human reader, and the structure that serves the human reader is the same structure that now serves the machine reader. The transformer is the first reader trained on both banks of the gap, the corpus of human language and the corpus of code, and therefore the first that can cross the bridge in both directions: it can read intent encoded as structure and emit code that encodes intent. Generative Specification works because it makes building and maintaining that bridge the primary act of development rather than a byproduct of it.
The two banks are not equal in this reader’s competence, and the inequality is the source of the bridge’s leverage. The training corpus is overwhelmingly natural language. Code is a small, exact, and fragmented slice of it, with each language its own precise grammar, several of them low-resource, and every token load-bearing in a way prose is not. The model is therefore substantially stronger on human conceptual meaning than on any specific language’s exact syntax. This asymmetry mirrors the human one, where a specification is easy to state and the exact code that satisfies it is where slips occur. The design consequence follows directly. Encoding intent in human-conceptual terms, through names, ubiquitous domain language, the specification, and contracts, routes the hard half of the problem, producing exact code, through the half the model is most fluent in, understanding what is meant. The bridge moves the load from the model’s weaker bank to its stronger one. This is why the derivation step is disproportionately determined by naming, by specification, and by the binding of tests to acceptance criteria: nearly all of that work is bridge-building, and the bridge is what makes a stateless reader’s output match human intent. Specifying in human-conceptual terms is therefore not a matter of convenience but the discipline’s central lever.
E. The Sentinel Navigational Tree
The derivability obligation introduces a failure mode of its own at scale, and the discipline must answer it structurally. If every intent must be externalized, a mature specification grows large. Loading the entire artifact set into a single session context reproduces, at the level of the context window, the incoherence the discipline exists to prevent. It also degrades accuracy directly: [Liu et al., 2023] demonstrate systematic accuracy degradation for information positioned away from the leading position of a long context window, so a reader given the whole specification at once attends least reliably to material buried in its middle, while consuming token budget on context the current task does not need. This subsection states the structural response: the sentinel navigational tree and a bounded working context.
The response is a sentinel navigational tree: a hierarchy of specification files in which each node declares its own scope and routes to its children. The root node is always loaded and must stay within a bounded line budget so that it occupies the leading, most reliably attended position. From the root the reader descends only the path relevant to the current task. The tree is lossless, meaning that joining all leaf nodes reconstructs the full specification, but each session receives only the slice it needs. This eliminates both the positional degradation and the unnecessary token expenditure of the monolithic load, while preserving completeness.
A well-formed tree must collectively contain five categories, enumerated in Table I. The first four are routinely present in mature documentation. The fifth, tool sequencing, is the category most often absent and the most consequential when missing. A specification that lists the tools available without stating when to prefer one over another, and in what order, forces the stateless reader to infer an operating sequence it has no reliable basis to infer, and inference at that point is exactly the arbitrary gap-filling the discipline exists to remove.
Table I. The Five Mandatory Categories of a Sentinel Navigational Tree
| Category | What it covers |
|---|---|
| Architectural identity | What the system is, its scope boundary, and the index of its architectural decisions |
| Standards | Naming, commit discipline, and quality-gate thresholds |
| Constraints and prohibitions | What must not happen, and the boundary violations the reader must refuse |
| Tool sequencing | When to use which tool, and in what order: not that the tools exist, but to use X before Y under condition C |
| Routing | What each child node covers and the condition under which the reader should descend into it |
Because each node declares its own scope, load policy, contributed categories, and routing targets, the tree is not merely a prose convention but a structure a gate can check by inspection: that the root stays within its line budget, that the five categories are collectively present across the nodes, and that every routing target resolves to an existing node. The tree is thus the concrete structural instrument by which the derivability obligation of Section III-B is satisfied without reintroducing the context-degradation failure the obligation would otherwise create at scale.