Developer labs¶
Use these labs to change one harness policy, predict the outcome, inspect the event sequence, and then reproduce the behavior locally. Each lab separates the interactive explanation from the evidence it can actually establish.
Start here¶
Recover from tool failure
Explore malformed calls, provider errors, lost acknowledgements, retry budgets, stable operation keys, and reconciliation.
Open the lab → 02Control the context budget
Compare replay, windows, protected evidence, retrieval, and compaction while watching omissions and integrity risks.
Open the explorer → 03Enforce approval across restart
Separate a reported pause from dispatch control, durable resume state, and replay-safe effects.
Open the clinic → 04Design a harness architecture
Turn workload, authority, restart, and delegation choices into a neutral component map and versioned dossier.
Open the builder → 05Match evidence to a claim
Compare versioned run records, expose incompatible controls, and keep observations separate from design guidance.
Open the workspace →How every lab works¶
- Predict what the harness will do before changing a control.
- Configure one policy without changing the workload.
- Inspect the timeline, effect count, and stopping reason.
- Run the executable fixture locally.
- Transfer the contract into your own harness or framework evaluation.
The browser interaction is a deterministic explanation of repository contracts. It is not a provider benchmark. The local commands are the executable evidence.
Continue with the tool failure and recovery lab, compose the contracts in the architecture builder, or return to Build your own harness.