01 / Research
Controlled experiments and investigations in LLM interpretability, alignment, agentic behavior, and more.
Recent examples
A security-first self-hosted agentic harness, with an integrated orchestrator, collaboration shell, sandbox, plugins, and keyless egress.
GitHub ↗
Safety-first multi-agent harness with dozens of tools, sophisticated memory management, and more.
Agent-agnostic shell: chat, kanban, terminal, and live browser takeover.
Dockerized coding agents, Chromium/CDP, desktop, and a full toolchain.
Credential-injecting egress so the agent process holds no real keys.
Jacobian lenses, sparse autoencoders, activation analysis, and controlled tests of what internal readouts establish.
Matched comparisons, behavioral probes, null baselines, leakage checks, and reproducible evaluation infrastructure.
Agent environments, retrieval systems, and evidence pipelines that connect model behavior to inspectable source material.
Code, prompts, data, hashes, artifacts, and result receipts ship with the claim.
Matched baselines and null tests determine what survives into the conclusion.
Negative results and broken hypotheses remain part of the public record.