praxagent / Methods
Experiment integrity
Plan comparisons before inspecting results. Report all outcomes, document changes, and publish the supporting evidence.
01 / Why
Credibility is chronological
Confirmatory research goes wrong when the rules change after the data arrive: optional stopping, selective reporting, or hypothesizing after results are known. We cannot erase judgment. We can make the consequential choices visible and time-ordered.
A dated record lets readers check that the design preceded the outcomes, the run followed the design, and all outcomes were retained. Later analyses should be labeled by when they were added.
02 / Sequence
Freeze, run, audit, release
Lock the design before outcomes exist
The question, claim boundary, sample, controls, endpoints, and analysis are committed before any target outcome is generated or inspected.
Execute the plan as frozen
Keep the planned rules independent of the observed effect. Document defects and necessary changes in dated amendments.
Bind the artifacts to the design
Runtime records are checked against the frozen plan, and every condition and failure path stays in the record.
Publish the supporting evidence
Code, prompts, data, and result records let readers inspect the evidence behind each claim.
03 / Labels
Say what the result is
These labels distinguish when the design was set and how the results can be interpreted:
Outcomes may still shape design
Useful for discovery and feasibility, and reported as exploratory rather than confirmatory evidence.
The design predates the outcomes
The complete plan was committed before target outcomes existed. Later changes are dated amendments, not quiet edits.
Everything reported was in the freeze
Endpoint, sample, controls, exclusions, and analysis all trace back to the frozen design.
Added after the results were seen
Sensitivity analyses can be informative, but they are labeled by timing and never promoted to confirmatory status.
04 / In public
What you should see
Release the evidence.
Published code, prompts, data, and result records let readers check the evidence behind each claim.
Run the control.
Matched baselines and null tests determine what survives into the conclusion.
Report the failure.
Negative results and broken hypotheses remain part of the public record.
These are our research standards. Each research note should identify its study plan, supporting results, and any changes made after outcomes were inspected.
A worked example: detecting SAE steering
The steering study links a protocol committed before outcomes, a report with results and evidence, and a dated follow-up amendment. The follow-up failed its planned replay-equivalence test, so its endpoint results are reported as exploratory. Readers can trace that change in status instead of inferring it from the headline.