Experiment publishing & validation — step by step
2026-09-28. For JC: exactly what to do with each experiment to make it count academically — where to pre-publish, how to make it credible, and whether/how to get someone independent to validate it. Built on the anti-monoculture plan (
experiments/revival/PREREGISTRATION.md, THREE-PAPER-ROADMAP §anti-monoculture). The honest core: pre-register before you run, publish a turnkey artifact with a DOI, and buy ONE axis of real independence — that is what converts a proponent’s demo into evidence.
The pipeline (per experiment, in order)
1. Pre-register — BEFORE running.
- Write the protocol (hypothesis, design, metrics, the null region, the named falsifier, the fit/holdout split). You already have the pattern in
experiments/revival/PREREGISTRATION.md. - Publish it to OSF (osf.io) — free, the standard registry, gives a timestamped, frozen record with a DOI. This is the anti-HARKing proof (you committed the design before seeing data).
- Commit the same file in the repo and log its git hash. OSF timestamp + git hash = double proof.
2. (Optional, stronger) Registered Report.
- Submit the protocol to a venue that accepts Registered Reports (ESEM, EMSE, and a growing set) for in-principle acceptance before you run. If accepted, the paper publishes whatever the result — null included. This is the single strongest answer to “proponent-authored,” because a neutral committee vetted the design before any data existed.
3. Run the experiment exactly as pre-registered. Deviations are allowed but must be disclosed as deviations (you already do this — the CR “weak-rung” deviation note is the model).
4. Publish the artifact (turnkey + citable).
- Package: benchmark + harness + oracle + raw data + the pre-registration + a README with step-by-step replication instructions (clone, setup, run, expected output).
- GitHub for the code (already there); tag a release and archive that release to Zenodo (GitHub↔Zenodo integration) → a DOI for the exact artifact version.
- If submitting to a venue with Artifact Evaluation (ICSE, FSE, ESEM), apply for the “Artifact Available / Evaluated / Reusable” badge. A badge is third-party confirmation the artifact runs — cheap credibility.
5. Pre-publish the paper.
- arXiv (cs.SE) for the preprint (citable immediately, establishes priority).
- Then submit to the target venue (empirical SE / IEEE / the RR track from step 2).
Independent validation — do you need it, who, and how
Yes — one axis of real independence is what the blind reviews demanded. Not all “validation” counts:
What does NOT count as independent (a reviewer prices these at zero):
- Graded capstone students you supervise · commercial partners (Gabriel/BYU) with a stake · the in-house LLM “stateless judge” running your own harness. These are your own incentive orbit.
What DOES count, and how to get it:
- A disinterested party runs the scoring/replication of the flagship claim. Recruit: another researcher (a co-author outside the project — e.g. the biologist collaborator for cross-discipline eyes; or a contact met through the advice-conversations), or pay a neutral contractor to run the turnkey artifact. One genuine independent run of the headline result is worth more than ten supervised ones.
- Call for replication on the published artifact (the benchmark-as-artifact strategy). Realistic expectation: a handful of downloads and a few “artifact available” citations; genuine adversarial re-runs are rare for a non-famous method. So treat the call as upside, not the plan — the independence in #1 is the plan.
- The human-rater study (DX2) — you must organize this one yourself, and it cannot be outsourced or automated away. To count: blinded raters, independently recruited (not your students), a pre-registered rubric, and a disinterested adjudicator for disagreements. This is the arm that breaks the “no human raters” critique for the rubric.
The honest boundary: the base paper’s experiments are already run and can ship with a full replication package now; independence for them is the call-for-replication + one disinterested re-run of the flagship (AX or KX). The owed experiments (D1 gating, D2 revival) get the full pipeline above, pre-registered from the start.
Where each thing lives (quick reference)
| Thing | Where | Gives you | |—|—|—| | Pre-registration | OSF (+ git hash) | timestamped frozen design (anti-HARKing) | | Design review before running | Registered Report venue (ESEM/EMSE) | in-principle acceptance; null publishes | | Code + harness | GitHub (tagged release) | the runnable artifact | | Citable artifact version | Zenodo (from the GitHub release) | a DOI for the exact data+code | | Preprint / paper | arXiv (cs.SE) → venue | priority + peer review | | Badge | venue Artifact Evaluation | third-party “it runs” |
The order to actually do it (this quarter)
- Base experiments (AX/KX/EX/RX/TX/SX/ALX/BX): assemble the replication package + Zenodo DOI, issue a call for replication, and line up one disinterested re-run of AX or KX. (They are already run; this is packaging, not new work.)
- Register the Compendium on Zenodo/arXiv as the citable master (the hub genspec.dev links to).
- D2 revival: it is already pre-registered — run it, publish the artifact, call for replication.
- D1 gating + the base-strengthening new experiments (2nd benchmark, capability ladder, DX2 human-rater): pre-register each on OSF before running.