Tool failure and recovery lab¶
Change the failure and recovery policy. The timeline shows which component acts, whether work is retried, and how many external effects remain. Predict the outcome first; then inspect why it happened.
Read the result as a harness contract¶
- Malformed calls and unknown tools belong on the model/tool boundary. Validate before dispatch, then return a correlated structured error so the model can correct the call.
- 429 and timeout failures may be retried, but attempts must share one total deadline and a bounded budget.
- 400 failures are invalid requests. Repetition amplifies load without changing the request.
- A timeout after commit creates an unknown outcome. A stable operation key prevents a second effect; independent reconciliation tells you what actually happened.
The interactive timeline is a deterministic simulation. It explains the policy but does not test a provider or framework.
Produce the same run record locally¶
python -m examples.harness.failure_lab --failure timeout_after_effect --retries 1 --stable-key --reconcile
The command emits versioned JSON with the selected policy, timeline, attempt count, effect count, and outcome. Try the unsafe variant:
python -m examples.harness.failure_lab --failure timeout_after_effect --retries 1 --no-stable-key --no-reconcile
Then run the process-crash fixture. It commits a synthetic booking, exits before checkpointing, restarts in a fresh interpreter, and proves through an independent SQLite sink that only one effect exists:
python -m examples.harness.reliability
pytest tests/test_failure_lab.py tests/test_learning_reliability.py -q
Compare framework behavior separately¶
The repository's resilience arena injects identical malformed model/tool calls into each adapter. That experiment measures whether the framework returns the error to the model and reaches the correction turn:
python -m arena run --arena resilience --framework all --mode mock --no-scorecard
python .github/scripts/check_resilience.py
Mock-mode differences are meaningful for this controlled recovery path because every adapter receives the same scripted fault. They still do not establish live-model quality. Read the published findings and run-mode boundaries.
Carry the contract into your harness¶
Record one terminal result per accepted tool call, stable operation identity for retryable effects, one total deadline, explicit retry classification, and an independent effect oracle. If any of these are missing, make the missing guarantee visible rather than inferring success from the final answer.
Next: Reliability and restart · Lost tool errors · Retry amplification