Context budget explorer¶
Choose a context policy and character budget. The explorer shows exactly which source records reach the model, what is omitted, and whether an obligation or correction disappeared. The fixed workload keeps the comparison about selection policy rather than model behavior.
Read the result as a harness contract¶
- Full transcript preserves the records but can exceed the request budget. A budget that is not enforced is only a dashboard value.
- Recent window is cheap and deterministic, but an early approval or constraint can disappear while the prompt still looks valid.
- Required + recent keeps explicit obligations first and fills remaining space with recent evidence. If required evidence cannot fit, it fails visibly.
- Required + retrieval limits candidates by a query and retains source IDs. Retrieval relevance does not establish authority or truth.
- Deterministic compaction demonstrates the space tradeoff without claiming semantic summary quality. The summary is labeled as a new record because its source-level provenance has collapsed.
This explorer budgets characters rather than provider tokens. Tokenization depends on the provider and encoding; use actual request usage for cost claims.
Produce the same selection locally¶
python -m examples.harness.context_lab --policy protected --budget 220
python -m examples.harness.context_lab --policy recent --budget 170
pytest tests/test_context_lab.py tests/test_learning_context.py -q
The command emits versioned JSON containing selected and omitted source IDs, missing obligations, used characters, outcome, and rendered context. The underlying context lesson also verifies durable memory isolation by tenant and session.
Compare growth separately¶
The repository's prompt-growth experiment measures bytes and estimated tokens across repeated turns. It answers how request cost grows; this lab answers which evidence a context builder chooses to retain.
python .github/scripts/report_growth.py 30
Read the prompt-growth evidence and the growing-context problem note. Mock token estimates are useful for controlled structure, not provider billing.
Next: Context, retrieval, and memory · Build your own harness