Skip to content

The controls the arena owns, and who checks they arrive

Why this page exists

Four fairness bugs of the same shape have been found in this repo, one at a time, each after a published number had already been built on it:

control what was actually happening found in
max_tool_iterations three adapters' "loop caps" were not loop caps — a budget of 6 ran 50, 41, and one too many #29-ish
request_timeout_s five of seven never passed it to their client and inherited a ten-minute library default transport.md
arena.tools (parameters) two adapters declared a narrower tool than the arena did, dropping a parameter the model never saw tool-schemas.md
arena.tools (descriptions) all six frameworks cut the sentence telling the model to pause — on the arena that grades pausing tool-schemas.md

Every one was invisible to mock mode, because the mock replays a script whatever the adapter sent. Every one would have shown up live as a difference in the framework's behaviour, and been attributed to the framework.

The pattern is not that adapters are careless. It is structural: a control the arena owns has to be carried by each framework in its own idiom — a different constructor argument, a different decorator, a different docstring convention — and nothing was checking the carrying.

So this page enumerates the controls instead of waiting for the next one to surface, and tests/test_shared_controls.py fails if a new one is added without an answer.

ArenaConfig

field does an adapter see it? held to it by
model yes tests/test_shared_controls.py
temperature yes tests/test_shared_controls.py
base_url yes implicit — nothing reaches the mock otherwise
api_key yes not gated; a wrong key fails loudly live (401) rather than silently
request_timeout_s yes tests/test_transport_faults.py
max_tool_iterations yes tests/test_adapters_contract.py
checkpoint_dir yes, durable arenas only tests/test_durable_across_a_restart.py
mode no — the harness starts the mock and rewrites base_url —
repeat no — the harness loops —
price_input_per_m, price_output_per_m no — the scorecard multiplies afterwards —

ArenaSpec

field held to it by
system_prompt tests/test_adapters_contract.py — the instruction must come from the arena, not the adapter
tools (which) tests/test_adapters_contract.py — no adapter may advertise a tool the arena did not declare
tools (shape) tests/test_tool_schema_fidelity.py — same parameters, same types, descriptions intact
dataset the harness iterates it; adapters never see it
durable tests/test_durable_state.py and tests/test_durable_across_a_restart.py

The gates on this page

Every ArenaConfig and ArenaSpec field is accounted for. Adding one now fails until someone has said which test holds adapters to it, or recorded that the harness consumes it and no adapter ever sees it. Cheap, and it is the half that would have caught all four bugs above at the moment the control was introduced rather than iterations later.

tools is worth noting: it is one field that went wrong twice, in different ways — first the set of tools an adapter advertised, then their shape. So it names both tests. "tools is checked" was true, and insufficient, the first time.

The configured model is the model on the wire. This is methodology §1 — one model for everyone, and no adapter picks its own — and it was the last fundamental rule with nothing behind it.

An adapter that quietly defaulted to its library's favourite model would be comparing a different model, which is the one difference that invalidates every number at once. It is also the one mock mode is least able to notice: the mock answers to any model name at all.

Measured across all seven: every adapter propagates it faithfully, including google_adk, which has to reach an OpenAI-compatible gateway through LiteLLM and so builds openai/<model> — LiteLLM strips the prefix and the name on the wire is exactly what the arena configured. A negative result, pinned because of what it would cost to be wrong about it later.

The configured sampling temperature is the one on the wire. methodology §3e — "temperature is 0.0 everywhere" was the last item in that list held by convention rather than by a field: ten hard-coded temperature=0.0 literals, one per adapter, checked by nothing. A framework left at its library default of 1.0 would sample differently from the rest of the field in every live run, for a reason that has nothing to do with the framework.

Measured across all fourteen buildable adapters: every one sends 0.0 today. ArenaConfig.temperature is now the single source, and the gate builds each adapter with a non-zero canary — a literal 0.0 left in an adapter sends the wrong number and fails.

No adapter puts a sampling parameter on the wire that the arena did not choose. Every request body is held to the envelope the hand-rolled baseline sends — model, messages, temperature, stream, tools, tool_choice — plus a per-adapter exception table. Measured: no adapter sets top_p, seed, max_tokens, the penalties, n, logprobs or response_format; smolagents alone adds stop (below). An adapter that started injecting top_p=0.9 would be comparing a different sampling distribution, and would fail here until the key was removed or recorded with a reason.

What is deliberately not gated

  • api_key. It reaches the adapter, and a wrong one fails loudly live with a 401 rather than silently skewing a comparison. There is no quiet failure to guard against.
  • smolagents' stop sequences. It sends stop: ["Observation:", "Calling tools:", …] on every request, in all three of its entries (smolagents, smolagents_multi, smolagents_code), because its text loops parse generation that halts at those markers — CodeAgent also stops at </code>. It is intrinsic to the mechanism, not a sampling choice, so it is an entry in the gate's exception table rather than a failure — the same treatment as strict: true and title on the tool schemas. The exception is itself checked to be live: tests/test_shared_controls.py fails if stop stops reaching the wire, so the waiver cannot quietly come to cover nothing while this page still describes it.
  • tool_choice. Every adapter sends "auto" today. A framework that drove its loop with "required" or "none" would be a mechanism difference worth recording, not a fairness break — the arena does not dictate how a framework decides to call a tool, only which tools exist.
  • Adapter-side prompt framing. Adapters may phrase the arena's instruction in whatever way is idiomatic — that difference is part of what is being compared. What they may not do is substitute their own task instruction, and that is gated.
  • Framework decoration of tool schemas. title, additionalProperties, strict: true and OpenAI's strict-mode widening of required are framework properties. Reported in tool-schemas.md, not failures.

Not measured

  • Mid-run drift, for the checks that only read the first request. Worth stating precisely rather than sweepingly, because the checks differ:

  • max_tool_iterations is aggregate — the probe counts every request that reaches a mock which never stops asking for tools, so an adapter that raised its own budget halfway through still fails.

  • model is aggregate too: the assertion is over the set of model names across every request, not the first.
  • tool schemas: the content checks read the first request only. Whether each parameter, type and description survives is asserted on the opening request. What is now checked on every request, for single-agent adapters, is that the arena-tool schema does not change after request 1 (tests/test_tool_schema_fidelity.py) — so an adapter that advertised the tools correctly and then narrowed them at request 3 is caught, even though the narrowed content itself is only inspected once. The _multi entries are excluded because they change the tool set on purpose when the speaker swaps, and a companion test asserts they really do, so the exclusion is not a blind spot.
  • request_timeout_s is measured on the first attempt's abandonment.
  • Live-mode-only controls. price_* never leaves the harness, so nothing here says the cost column is right — only that the token counts feeding it are (usage accounting).