praxagent / Methods
Experiment integrity
Freeze the design before outcomes. Run what you froze. Keep the failures. Ship the receipts with the claim.
01 / Why
Credibility is chronological
Confirmatory research goes wrong when the rules change after the data arrive: optional stopping, selective reporting, or hypothesizing after results are known. We cannot erase judgment. We can make the consequential choices visible and time-ordered.
The point is not ceremony. It is a public trail that the design preceded the outcome, the run followed the design, all outcomes were retained, and later analyses are labeled by when they entered the story.
02 / Sequence
Freeze, run, audit, release
Lock the design before outcomes exist
The question, claim boundary, sample, controls, endpoints, and analysis are committed before any target outcome is generated or inspected.
Execute the plan as frozen
No editing rules mid-flight because an effect looks weak, strong, or inconvenient. Defects become dated amendments, not quiet fixes.
Bind the artifacts to the design
Runtime records are checked against the frozen plan, and every condition and failure path stays in the record.
Ship the receipts with the claim
Code, prompts, hashes, artifacts, and result receipts are published so a reader can inspect what the claim rests on.
03 / Labels
Say what the result is
Not every run is confirmatory. We keep status language boring and consistent:
Outcomes may still shape design
Useful for discovery and feasibility, and labeled as such — never sold as a confirmatory error rate.
The design predates the outcomes
The complete plan was committed before target outcomes existed. Later changes are dated amendments, not quiet edits.
Everything reported was in the freeze
Endpoint, sample, controls, exclusions, and analysis all trace back to the frozen design.
Added after the results were seen
Sensitivity analyses can be informative, but they are labeled by timing and never promoted to confirmatory status.
04 / In public
What you should see
Release the evidence.
Code, prompts, data, hashes, artifacts, and result receipts ship with the claim.
Run the control.
Matched baselines and null tests determine what survives into the conclusion.
Report the failure.
Negative results and broken hypotheses remain part of the public record.
This page is the public posture, not the internal checklist. Research notes carry the concrete receipts for each claim.