Quality gates
A specification is only as strong as what a machine can verify. Gates are the deterministic checkers that make the method enforceable: standard CI hooks, grouped here by the rubric property each one defends. An LLM may write the code, but a non-LLM checker verifies it. These are those checkers. They are the verification and enforcement layer of the substrate (the harness), the part that keeps the guarantee outside the model.
Terminology: guides and sensors. Böckeler (2026) divides an agent’s harness into guides (feedforward) and sensors (feedback), each computational or inferential. Here the sentinel, specifications, instruction files and skills are guides; tests, linters, structural checks and gates are sensors, and “harness” used for the verification layer means the sensors. An instruction file is advisory, a hook is deterministic; both layers should derive from the same ratified specification, with the guides kept minimal. Böckeler, B. (2026). Harness engineering. martinfowler.com. https://martinfowler.com/articles/harness-engineering.html
You do not need a proprietary tool. Point a capable assistant at the white paper and this page, and ask it to wire the gates that apply to your stack. The tools below are the JavaScript/TypeScript defaults; the assistant substitutes per language.
Blocking versus advisory. A blocking gate fails the build. An advisory gate reports and does not stop the merge. A gate that can be skipped under deadline pressure is a suggestion, which is the failure the rubric calls Open Gates. Advisory rows below are the ones where a hard threshold is still project-specific.
The gates, by property
The library column names real gates in quality-gates/, where one exists. A dash means the row is a practice with no library entry yet.
Bounded
| Gate | Standard tool | Rule | Level | Library gate |
|---|---|---|---|---|
| File length | wc / loc | at most about 300 lines | blocking | file-length-max-300 |
| Function length and parameters | eslint | max-lines-per-function, max-params | blocking | function-length-max-50, max-function-parameters |
| Cyclomatic complexity | eslint complexity | function complexity at most 10 | blocking | cyclomatic-complexity-max-10 |
| Code duplication | jscpd | no new duplication in the diff above a stored baseline (ratchet) | advisory | no-duplicated-code-in-diff |
| Dead code | ts-prune / knip | no unused exports | advisory | no-unused-exports-dead-code |
Composable
| Gate | Standard tool | Rule | Level | Library gate |
|---|---|---|---|---|
| Layer boundaries | dependency-cruiser | controller does not reach the repository; domain stays pure | blocking | no-direct-db-in-routes |
| Circular imports | madge | no cycles | blocking | no-circular-dependencies |
Verifiable
| Gate | Standard tool | Rule | Level | Library gate |
|---|---|---|---|---|
| Type strictness | tsc --strict | strict passes; no any | blocking | typescript-strict-mode, no-any-type |
| Line and branch coverage | c8 / istanbul | at or above a target | blocking | coverage-threshold-80 |
| Mutation score | stryker | at or above a target (the library uses 65% overall, 70% on changed files) | advisory on the site’s reference set; blocking in the library | mutation-score-threshold |
Coverage and mutation score are complementary, not interchangeable. Coverage measures execution; mutation score measures detection. In the AX study the first GS treatment reported 93.1% line coverage but a 58.62% mutation score; after three rounds of assertion improvements the mutation score converged to 93.10%. Both gates are needed.
Executable
| Gate | Standard tool | Rule | Level | Library gate |
|---|---|---|---|---|
| Behavioural probes against a live system | hurl (or newman) | acceptance criteria pass on deploy | blocking | hurl-contract-tests-pass, contract-tests-against-live-env |
| Module boot smoke | CI boot script | every module boots | blocking | smoke-test-passes, health-endpoint-responds |
Executable is scored only when a formal behavioral contract exists to run; the rubric page has the rule. Related library gates that the library files under this property: jest-no-failed-tests, tsc-no-emit-exits-zero, docker-compose-defined.
Defended
| Gate | Standard tool | Rule | Level | Library gate |
|---|---|---|---|---|
| Secrets scan | gitleaks | no secrets in the diff | blocking | no-hardcoded-secrets |
| Dependency audit | npm audit / osv | no high or critical vulnerabilities | blocking | npm-audit-no-high-cve, container-image-no-critical-cve |
| Forbidden patterns | eslint custom rules | for example no eval, no database access in controllers | blocking | no-debug-routes-in-production, tls-enforced |
A dependency audit and an architecture audit are independent checks. In AX, Treatment-v2, the first condition to reach 12/12 on the rubric, also carried nine high-severity vulnerabilities against zero for the control, from a dev-dependency chain; one prescriptive directive plus an npm audit gate took it to zero (single model, single benchmark). Neither check subsumes the other.
Defended has a human ceiling. CI can verify six of the seven properties automatically; whether adversarial challenge has been anticipated needs human review, so an automated Defended score is provisional. CI runners are also external infrastructure that generated code cannot itself provision, which is why a spec must name the hook and gate files to be emitted rather than merely describe them.
Self-describing
| Gate | Standard tool | Rule | Level | Library gate |
|---|---|---|---|---|
| Sentinel present, names announce the domain | structure check | the root sentinel routes to the spec slices | advisory | none |
| Setup and configuration documented | file / section check | a runnable setup section; every environment variable documented | blocking | readme-setup-section, env-vars-documented, jsdoc-public-functions |
Auditable
| Gate | Standard tool | Rule | Level | Library gate |
|---|---|---|---|---|
| Conventional commits | commitlint | commits parse by type and scope | blocking | conventional-commits |
| Decision records | commit-history and file hook | every referenced ADR exists as a committed file with content | blocking | adr-files-emitted |
| TDD phase order | commit-history hook | test:[RED] before feat: | advisory | none |
When to run each check, and how to remediate a finding without weakening the gate: Run structural gates and remediate.
The ratchet
Each new defect that slips through becomes a new blocking gate, derived from a real incident. The set only grows and only tightens. In the AX series, each gap a run exposed was closed as a template or gate change (mutation gate, emit-don’t-reference, dependency governance, a DRY gate) that then applied to every governed project. That accumulation, not any single gate, is the value.
The community library
The quality-gates/ directory is an open library of structured gates, each mapped to a GS property, with a schema and a contribution path. The gate files are the source of truth: the library page is generated from them, so it is always the current list, and this page deliberately states no count. The library also holds gates for academic papers (for example claim-scope-calibration, notation-audit), which apply the same idea to a document instead of code.
A note on classification: the library files some gates under a different property than the grouping above (for example no-any-type under Bounded, typescript-strict-mode under Executable). The grouping here follows what each gate defends in the method; the library follows its schema. Composable is the least represented property and the highest-value place to contribute.
Browse the library How to contribute a gate
What gates do not claim
- Gates verify. They do not make a specification correct; a gate only checks the obligations someone wrote down. The completeness of the spec is a separate question, covered in spec completeness.
- Raising a score by the letter of a check, without the underlying property, is the failure the rubric exists to catch.
- No claim is made here about how many gates a project needs. The evidence page says what has been measured and what has not.
Next: The evidence · The rubric · Spec completeness