Evidence workspace¶
Start with a claim, then compare two run manifests. The workspace tells you whether the evidence mode can support the claim and whether both runs share the controls required for a direct comparison.
Inspect the versioned records
Read the verdict carefully¶
- Mock establishes wiring and mechanics under fixed responses.
- Offline fixture establishes a local contract under controlled faults and an independent oracle.
- Codex functional exercises a real model through the subscription bridge. Translation prevents native latency, token, and cost comparisons.
- Live provider can support provider and framework comparisons when the model, task, tools, configuration, scorer, and exclusions are held constant and repeated.
“Comparable” means the manifests describe the same evidence contract. It does not mean the result is statistically stable or the causal explanation is known. A provider comparison with fewer than three repetitions remains visibly weak even when its fields match.
Create a run record locally¶
python -m examples.harness.evidence_lab --claim recovery --mode offline-fixture --observation "one effect after retry"
python -m examples.harness.evidence_lab --claim provider_comparison --mode live --model MODEL --repetitions 5
pytest tests/test_evidence_lab.py -q
The record uses agentic-arena.run-record/v1 and keeps observations separate from design_guidance. It records commit, configuration, model, dataset, scorer, repetitions, and exclusions. The browser download adds both records and their comparison under a versioned envelope; its share link encodes the choices, not unpublished results.
Use run modes to choose evidence and methodology before publishing a comparison.
Next: Developer labs · Evaluate a claim