Work Research Notes Knowledge Base Search Contact

praxagent / Methods

Experiment integrity

Freeze the design before outcomes. Run what you froze. Keep the failures. Ship the receipts with the claim.

01 / Why

Credibility is chronological

Confirmatory research goes wrong when the rules change after the data arrive: optional stopping, selective reporting, or hypothesizing after results are known. We cannot erase judgment. We can make the consequential choices visible and time-ordered.

The point is not ceremony. It is a public trail that the design preceded the outcome, the run followed the design, all outcomes were retained, and later analyses are labeled by when they entered the story.

02 / Sequence

Freeze, run, audit, release

01 / Freeze

Lock the design before outcomes exist

The question, claim boundary, sample, controls, endpoints, and analysis are committed before any target outcome is generated or inspected.

02 / Run

Execute the plan as frozen

No editing rules mid-flight because an effect looks weak, strong, or inconvenient. Defects become dated amendments, not quiet fixes.

03 / Audit

Bind the artifacts to the design

Runtime records are checked against the frozen plan, and every condition and failure path stays in the record.

04 / Release

Ship the receipts with the claim

Code, prompts, hashes, artifacts, and result receipts are published so a reader can inspect what the claim rests on.

03 / Labels

Say what the result is

Not every run is confirmatory. We keep status language boring and consistent:

Exploratory

Outcomes may still shape design

Useful for discovery and feasibility, and labeled as such — never sold as a confirmatory error rate.

Frozen

The design predates the outcomes

The complete plan was committed before target outcomes existed. Later changes are dated amendments, not quiet edits.

Confirmatory

Everything reported was in the freeze

Endpoint, sample, controls, exclusions, and analysis all trace back to the frozen design.

Post-run

Added after the results were seen

Sensitivity analyses can be informative, but they are labeled by timing and never promoted to confirmatory status.

04 / In public

What you should see

01

Release the evidence.

Code, prompts, data, hashes, artifacts, and result receipts ship with the claim.

02

Run the control.

Matched baselines and null tests determine what survives into the conclusion.

03

Report the failure.

Negative results and broken hypotheses remain part of the public record.

This page is the public posture, not the internal checklist. Research notes carry the concrete receipts for each claim.

Research notes Back to How we work