Skip to content

Findings

Everything this benchmark has actually measured, in one place, with the command that regenerates each number.

The per-framework pages and the topic docs are where the reasoning lives; this page is the index of results. If a claim is not on this page, it is not something the repo has measured.

How to read it

Three rules govern what appears here, and they are the whole reason the numbers are worth anything:

  1. Everything that could vary is held identical. Same model, same gateway, same tools, same task spec, same eval set, same iteration budget, and — in mock mode — byte-identical scripted turns. See methodology.md.
  2. Only two things compare honestly offline, and every finding below is one of them or a direct wire measurement: what a framework puts on the wire, and how it behaves when the model misbehaves. Mock-mode pass rate is ~100% by construction and means nothing. (methodology §5)
  3. Differences between frameworks are findings, not test failures. They are reported, not gated. What CI gates are the invariants — that adapters report their costs honestly, respect the budget, and pass the tool result through unaltered.

Every number below was regenerated by CI on a clean Linux install, not only locally.


1. What a framework costs you on the wire

Six of seven frameworks sit inside a 1.15× band; one sits at 3.90×. Mean prompt tokens per item on tool_use, 15 items:

framework prompt tokens vs baseline
vanilla (stdlib baseline) 753.5 1.00×
langgraph 753.5 1.00×
pydantic_ai 794.0 1.05×
microsoft_af 802.0 1.06×
google_adk 836.1 1.11×
openai_agents 856.9 1.14×
smolagents 2935.5 3.90×

No framework is leaner on the wire than the hand-rolled loop — and this page said the opposite until it was checked properly. See §5: the frameworks that measured cheaper were sending less, and with the tool schemas equalised langgraph ties the baseline byte for byte. Within the band the entire spread is still the tools block — the messages are byte-identical at 472 characters — but what separates the others from the floor is decoration (title, additionalProperties, strict: true), not tighter serialisation.

smolagents' 3.90× is a system prompt, not tool serialisation. A 4,207-char template goes out on every request where the arena asked for 384, and 1,159 of those characters are a prose restatement of the same tools already sent as an OpenAI schema — the tools are transmitted twice and billed twice. It is doing real work (driving models that cannot natively tool-call), but paired with a model that tool-calls well it is ~3.8 KB per request of scaffolding you do not need.

Its CodeAgent shape is heavier again — 6.95×. smolagents_code runs the same tools through a Python-writing loop instead of native tool calls; a longer few-shot system prompt puts it well above the ToolCallingAgent entry, though its completions run terser (60.7 tokens against 74.7). Same 2.07 LLM calls. The side-by-side table is on the smolagents page.

→ overhead.md · tool-schemas.md · smolagents.md

python -m arena run --arena tool_use --framework all --mode mock --no-scorecard && python .github/scripts/report_overhead.py tool_use

1b. …and what happens to that cost when the loop gets long

The table above is a two-call task, which is where framework overhead looks its worst. Running the same conversation out to 30 tool-calling turns:

request # 1 11 31
vanilla 121 1531 4350
smolagents 1069 2551 5515
ratio 8.83× 1.67× 1.27×

Nobody truncates. Not one of the seven drops, windows or summarises anything; all 31 requests carry the entire history. None of these libraries ships context management on the default path, so the only thing bounding your prompt in a long loop is max_tool_iterations — the harness's knob, not the library's.

Frameworks differ by a constant, not by a rate. One more turn costs everyone 136.7–148.2 tokens, an 8% band, against an 8.8× spread on the first request.

Those two facts explain the decay: a per-request constant is linear in turns, the conversation it is divided by is quadratic, so any fixed overhead decays as 1/n. vanilla spends 9,901 estimated prompt tokens on a 10-turn item and 71,745 on a 30-turn one — 3× the turns, 7× the bill. Read 3.90× as an upper bound on short items, not a tax on every item.

→ overhead.md

python .github/scripts/report_growth.py 30

2. What breaks when the model misbehaves

Eight scripted faults, byte-identical for every framework, so any difference is the framework's own error handling. This is the other thing mock mode compares honestly.

framework recovered fails on / at what cost
vanilla 8/8 — (2 LLM calls each)
pydantic_ai 8/8 —
microsoft_af 8/8 —
langgraph 7/8 res-01 — malformed tool arguments
openai_agents 7/8 res-02 — raises on an unknown tool name
google_adk 6/8 res-01 and res-02, both uncaught exceptions
smolagents 8/8 but 6 LLM calls (the whole budget) on the four faults its validator rejects before the tool runs, against 2 elsewhere

smolagents' rough edge is structural, and now shows up as cost rather than a lost item. The split is exact: on the four faults where the tool ran and returned something it recovers in two calls like everyone else; on the four its validation layer rejects first (unknown name, missing argument, unexpected argument, null arguments) the error never reaches the transcript, the model cannot see it, and it re-emits the identical call until the step budget is gone — six calls and ~2.7× the prompt tokens, with zero tool calls recorded. It still produces the right final answer at the end, so the item passes; a retry wrapper does not help, because the prompt is byte-identical each time.

Correction. This row read 4/8 for several iterations. That was measured when an exhausted smolagents run returned ""; it now surfaces a final answer from its last memory step instead, so numeric_equals passes. The failure mechanism is unchanged — the validator-rejected faults still burn the whole budget with no tool call — but the consequence is a 3× cost, not a lost item. check_resilience.py now gates every count in this table so the next drift fails CI instead of ageing silently.

Google ADK is the only framework that loses both res-01 and res-02, and both are uncaught exceptions rather than the model giving up. Loud, at least, and the kind of thing a retry wrapper does handle.

→ decision-guide.md §2 · feature-matrix.md

python -m arena run --arena resilience --framework all --mode mock --no-scorecard && python .github/scripts/check_resilience.py

2b. A quieter failure: what happens to a batch containing one broken call

When a model returns several tool calls in one turn and one of them is broken, two questions of the next request — did the successful call's result reach the model, and was the broken call reported at all?

unknown tool malformed args missing required arg
vanilla, pydantic_ai, microsoft_af both both both
langgraph both good only both
smolagents error only good only error only
openai_agents raises both both
google_adk raises raises both

Three ways to mishandle a batch, and they are not equally bad. Silent partial — the broken call vanishes with no message of any kind, and the run answers from partial evidence without the model ever being told. The successful sibling discarded — smolagents surfaces the error, but as a rewritten task rather than a turn in the transcript, dropping the good result with the history, so the model re-emits the identical batch. Raises — loud, and the best of the three, because nothing is silently wrong.

This refines the res-01 result above. LangGraph losing malformed arguments is conditional: alone, the call produces no tool message and the graph halts, so the item fails visibly; batched with a call that succeeds, the graph carries on and the drop becomes invisible. The visible failure is the better outcome.

→ feature-matrix.md

python -m pytest tests/test_parallel_tool_calls.py -q

2c. When the gateway fails, not the model

Every arena scripts the model. None scripts the provider returning 429, 500 or 400 — which is what real deployments hit. Attempts that reached the wire:

framework 429 once 429 ×3 500 once 400
vanilla (stdlib baseline) raises (1) raises (1) raises (1) raises (1)
langgraph ok (2) raises (2) ok (2) raises (1)
pydantic_ai, openai_agents, microsoft_af, google_adk ok (2) raises (3) ok (2) raises (1)
smolagents ok (2) ok (4, +2–4 min) ok (2) raises (1)

This is the clearest answer in the repo to "what does a framework buy you?" The hand-rolled loop has no retry at all: one 429 and the item is lost. Every framework survives it. It is the first dimension where the baseline loses outright — on prompt size it is mid-pack, on resilience joint best, on delegation exactly as expensive as a graph library.

smolagents is the only one that survives three consecutive 429s — by sleeping for minutes on one item (measured five times at 139–225 s, so heavily jittered). Nothing in the scorecard shows it: the item passes. In a batch that is throughput quietly collapsing. Everyone else fails fast in ~1.3 s and hands you the error. Both are defensible; neither is documented.

LangGraph retries once where the others retry twice, and everyone correctly refuses to retry a 400.

Every framework that retries honours Retry-After exactly — 3.01 s for a header of 3, 9.02 s for 9. But smolagents' long sleep is not shortened by it: there are two retry layers, and only the inner one (the OpenAI client) listens. The outer one is smolagents.models.Retrying, whose first backoff is RETRY_WAIT(60) × 2 × (1 + random()) — 120 to 240 s, bracketing every measurement, and it never consults the header.

But that minutes-long sleep is rate-limits only. Against a provider that simply hangs, smolagents fails in 8.5 s like everyone else — its outer layer is built with retry_predicate=is_rate_limit_error, a substring check for 429 / rate limit in the exception text, and a timeout matches none of it.

A hung provider also caught a bug in this harness, not in the frameworks. ArenaConfig.request_timeout_s has existed since the first commit, and five of the seven adapters never passed it to their client — pydantic_ai, openai_agents, microsoft_af, smolagents and google_adk each inherited their library's default (ten minutes, on the official OpenAI client) and answered happily through a 20-second hang on a one-second budget. Fixed, and gated twice: that a hang is abandoned, and that doubling the budget doubles the wait, which is what stops any hard-coded value passing.

The knob turns out to bound an attempt, not an item. A framework that retries twice spends timeout × attempts + backoff, so the default 60 s is a three-minute worst case per item.

→ transport.md

python .github/scripts/report_transport.py

3. What delegation costs

multi_agent carries real three-role pipelines (researcher → writer → editor) alongside the single-agent entries, with identical role wording, so three questions come apart:

comparison prompt LLM calls
vanilla → vanilla_multi — three stages, framework-free 2.50× 2.00×
langgraph → langgraph_multi — the same, inside a framework 2.50× 2.00×
vanilla_multi → langgraph_multi — the graph machinery alone 1.00× 1.00×
openai_agents → openai_agents_multi — native handoffs 2.64× 2.00×
smolagents → smolagents_multi — sub-agent invoked as a tool 3.93× 3.00×
pydantic_ai → pydantic_ai_multi — the same, hand-built 3.57× 3.00×

The cost of multi-agent is the structure, not the framework. Three roles double the LLM calls and 2.5× the prompt tokens, and cost that whether you build them with a graph library or a for loop.

The graph machinery adds nothing — vanilla_multi and langgraph_multi send the byte-identical request. langgraph used to carry a 48.6-token tool-schema difference from vanilla (a 0.97× that read like "the framework is cheaper"); a LangChain release equalised the serialisation, so it now matches byte for byte here as well as on tool_use.

Correction. The row above read langgraph_multi 2.62× and the graph machinery 0.97×, and openai_agents_multi 2.76× / smolagents_multi 4.03×, for several iterations. A LangChain / openai-agents / smolagents release bump moved each one's serialisation; vanilla and pydantic_ai are unchanged, every call multiplier is unchanged, and report_delegation.py now gates all five ratios.

You pay for a handoff by advertising it, not by taking it. Model-decided delegation costs ~10% more prompt than the structural pipeline at the same call count. Decomposing that gap across one item's four requests:

messages tool schemas
vanilla_multi 6767 722
openai_agents_multi 6818 1495
difference +51 +773

The transfer call and its result add 51 characters to the entire conversation. The transfer_to_* schemas add 773 — 94% of the difference — because they ride on every request whether or not anyone delegates. A supervisor offering N handoffs pays for N schemas on every request in the run: the cost scales with the options you offer, not the ones you use.

Three roles cost 2× the calls in every mechanism except one. A sub-agent invoked as a tool costs 3×, because its reply is a tool result rather than the end of the run: the manager is still running and has to produce its own final answer afterwards. Visible on the wire as six requests, down the chain and back up it — the last two are pure mechanism. So this cost grows with the depth of a chain rather than with the work in it. (The 3.93× is a floor: the mock forwards only the task, so the writer never receives the researcher's findings. With a speaker swap the transcript comes along free; with a sub-agent you pass context by hand and pay for it again.)

And it generalises to a second, structurally different mechanism — at 3.3× the price. smolagents' managed_agents advertises a sub-agent as an ordinary tool rather than as a transfer that swaps the speaker. Swept the same way as the handoff, both curves are linear in delegates offered: managed_agents costs ~875 characters per delegate on every request (after a ~385-char one-off preamble), the OpenAI Agents SDK's handoffs ~251, from the first one, with no preamble and nothing added to the system prompt. The gap is the same one behind smolagents' 3.90× in §1 — each managed_agent is described twice, as a JSON schema and again in prose, while a transfer_to_<name> schema is a name-templated stub with an empty parameter object and no description of the target at all. Both frameworks bill you for options rather than actions; they do not bill the same amount, or in the same place.

Measured from one role to five, across five implementations in four libraries, every one follows an exact law:

roles 1 2 3 4 5
handoffs (openai_agents) speaker swap 2 3 4 5 6 N + 1
sub_agents (google_adk) transfer, returns to parent 2 4 5 6 7 N + 2
managed_agents (smolagents) sub-agent as a tool 2 4 6 8 10 2N
AgentTool (google_adk) sub-agent as a tool 2 4 6 8 10 2N
agent delegation (pydantic_ai) sub-agent as a tool 2 4 6 8 10 2N

2N is the mechanism, not the library — three implementations in three libraries that share no code, agreeing at every depth. The third one settles it: Pydantic AI has no delegation feature, so its chain is an ordinary async tool whose body awaits a nested run, and it costs exactly 2N anyway. The price is in the shape — a nested run returns a value instead of ending the conversation — and no library choice avoids it. "Handoff" is not one thing — both the OpenAI SDK and ADK call theirs a transfer, and ADK's costs one call more because it returns control to the parent, which then speaks again.

And you pay in calls or in prompt, not both — the call law alone points the wrong way. Inside ADK at four roles: sub_agents costs 6 calls and 7375 prompt tokens; AgentTool costs 8 calls and 1722. A transfer keeps one conversation every agent sees all of, so the prompt compounds (15.4× against 3× the calls); a sub-agent starts fresh, so the prompt stays flat and the calls compound instead (3.59× against 4×). Cheaper in calls is dearer in prompt, and prompt is usually the larger bill.

Restarting only helps if there is little to re-send — and the third sub-agent-as-tool implementation is what proves that is the library rather than the mechanism. smolagents restarts each sub-agent fresh and still compounds (5.68×); Pydantic AI restarts them exactly the same way and lands on ADK's numbers instead — 1581 prompt tokens at four roles against ADK's 1722 and smolagents' 13441. So smolagents' outlier is its own ~4 KB templated system prompt, re-sent by every fresh sub-agent (the scaffolding from §1), and nothing about starting fresh. Note too which implementation is cheapest in prompt at four roles: the one that costs the most calls.

Forwarding context does not change the ranking — it widens the gap. The obvious objection to everything above is that every pipeline here is measured with the delegate told only the task, which favours whichever mechanisms carry the transcript for free. Handing the same payload to every framework, extra prompt tokens per item against forwarding nothing:

forwarded characters 0 277 553 1105
vanilla_multi, langgraph_multi — structural 0 0 0 0
openai_agents_multi — speaker swap 0 0 0 0
smolagents_multi — sub-agent as tool 0 509 978 1916
pydantic_ai_multi — sub-agent as tool 0 508 977 1915

Three of the four mechanisms forward for free — a structural pipeline shares one transcript, and a speaker swap's transfer_to_<agent> tool declares no parameters at all, so there is nowhere to put a payload. Only a sub-agent starting a fresh conversation has to be told, and the two implementations of that agree to within one token at every size, in libraries that share no code.

It costs about 1.77 tokens per forwarded character, not the 0.25 of a single copy: the payload becomes the sub-agent's opening message and rides every subsequent request of its conversation. So 3.93× and 3.57× were floors, and a modest 553-character findings block takes pydantic_ai_multi to roughly 4.9× while openai_agents_multi stays at 2.64×. The gap about doubles; it does not invert.

Two things this deliberately does not say. The lower completion tokens on the handoff chain (140 vs 198) are not efficiency — the researcher emits a short transfer instead of a draft, so only the last agent writes a brief. And the mock scripts identical turns for all four entries, so the benefit is held at zero by construction: this is the bill delegation puts on the table before it has done anything for you.

→ multi-agent.md

python -m arena run --arena multi_agent --framework all --mode mock --no-scorecard && python .github/scripts/report_delegation.py

4. Pausing, and surviving a crash

Six frameworks pause; no two of the six mechanisms look alike. human_in_the_loop observes the pause in the harness rather than trusting the agent's prose, and durable_state throws the runner away at the pause and rebuilds it, so only a real checkpoint gets across.

adapter pause crash mechanism
langgraph 12/12 8/8 interrupt() + MemorySaver / SqliteSaver, resumed with Command(resume=…)
openai_agents 12/12 8/8 needs_approval=True → ToolApprovalItem; RunState.to_json() serialises the whole run
pydantic_ai 12/12 8/8 tool raises CallDeferred → DeferredToolRequests; stateless replay as message_history
google_adk 12/12 🟡 8/8 LongRunningFunctionTool; DatabaseSessionService on sqlite+aiosqlite
microsoft_af 12/12 ❌ approval_mode="always_require" + ToolApprovalMiddleware
vanilla 12/12 🟡 8/8 emulated — transcript carried back in; durable, but not a checkpointer

Google ADK reports the pause but does not enforce it. Left alone, ADK hands the model the {"status": "pending"} marker and the model carries on — measured, it went straight to book_room with no human decision, which is exactly what the arena's no_tool_before_suspend check exists to catch. Breaking out of the event stream is what makes the pause real. It is the only one of the six where ignoring the signal lets the agent act, hence 🟡.

microsoft_af's durable gap was measured, not assumed. Its approval state serialises cleanly (~400 chars of clean JSON), but the conversation lives in the AgentSession's in-memory store, which returns raw strings after a JSON round trip, and restoring the approval state into a rebuilt agent re-queues the request instead of consuming the answer. The adapter therefore does not expose resume on a durable arena at all — keeping one it cannot honour would post 0/8 and read as a broken framework rather than an absent feature.

The durable_state check is not vacuous. An adapter patched to restart from scratch instead of resuming drops to 0/8: it still reaches the right answer, by redoing both lookups, and the call_counts check catches the duplicated work.

What the pause costs, and why "fits in a message list" is the whole point. resume_state is the only thing carried across the restart. On the two-lookup durable_state item:

adapter resume_state on-disk store what is in the state
google_adk 112 B 57 KB sqlite a session id
langgraph 401 B 76 KB sqlite a thread id + config
vanilla 2.5 KB — the whole transcript
pydantic_ai 6.1 KB — message history as JSON
openai_agents 17 KB — RunState.to_json() — the entire run object

A checkpointer's resume_state is a handle — a few hundred bytes whatever the conversation did — because the run lives in its store (a fixed sqlite floor, not proportional to turns). A serialiser's resume_state is the conversation, so it grows with every turn, and openai_agents is heaviest because it serialises the run object rather than just the messages. Both survive the restart; only the first kind still does once the transcript is large. The under-1-KB handle and the several-KB transcript are gated as a split in tests/test_durable_across_a_restart.py so a checkpointer that started inlining the transcript would fail.

→ feature-matrix.md · per-framework pages in frameworks/

python -m arena run --arena human_in_the_loop --framework all --mode mock --no-scorecard
python -m arena run --arena durable_state --framework all --mode mock --no-scorecard

5. Can the numbers themselves be trusted?

Two findings that are about the benchmark rather than about any framework. Both came from asking what nothing was checking.

durable_state could be passed by cheating, and now cannot. The arena's claim is that the runner is thrown away at the pause. The harness enforced that by rebuilding the runner and JSON round-tripping the state — both inside one interpreter, which a module-level cache keyed by item id survives. Measured rather than supposed: a wrapper doing exactly that scores 8/8, with ['search', 'search', 'calculator'] and one suspend, indistinguishable in the scorecard from the honest baseline it wraps.

Durability now runs its two legs in two processes, with nothing between them but a JSON file and checkpoint_dir. All five resumable adapters pass — a negative result, and the reason the published 8/8s mean what they say — while the cheat raises KeyError. The cheat ships as a test fixture so a green file cannot come to mean "the probe ran and asserted nothing".

It also settles the stateless-resume / real-checkpointer split by measurement rather than by reading the source: langgraph writes langgraph.sqlite and google_adk writes adk_sessions.sqlite; vanilla, pydantic_ai and openai_agents write nothing and carry the whole conversation in resume_state. → methodology.md §7

Nothing was verifying that adapters report cost honestly — every published comparison on this page assumed it. A gate holding each AgentResult against what the mock server actually served found a real bug immediately: langgraph reported 3 LLM calls on durable_state where 4 were served, dropping a whole leg. Correctness was never affected (8/8 before and after), which is precisely why it survived. The same gate then caught a second, different accounting bug in the very next adapter written against it. Every other adapter verified honest — a negative result, and the thing that makes §1 and §3 worth reading.

The frameworks were not describing the same tool. §1 said for fifteen iterations that the hand-rolled baseline is not the leanest — that three of four frameworks serialise the same tools more compactly. Nothing had compared what those bytes say, only what they cost. Against the canonical spec in arena/tools/__init__.py:

  • google_adk and smolagents declared search(query) where the arena declares search(query, k=3) — a parameter the model was never offered;
  • langgraph, pydantic_ai, openai_agents and microsoft_af sent bare types, no parameter descriptions, because none of these libraries reads one from the docstring and each wants it somewhere different.

Mock mode is structurally blind to this: scripted calls are replayed whatever the schema said, so every arena stayed green. Live it would have read as a quality difference and been blamed on the framework.

With the tools equalised, langgraph ties the baseline exactly — 753.5 tokens against 753.5, tools block 637 chars against 637. Its whole 0.91× advantage was four missing descriptions, and the corrected claim is the duller one: no framework is leaner than the hand-rolled loop, and the band is 1.00×–1.14× rather than 0.91×–1.05×. The second correction to that claim, having already replaced an earlier hypothesis that the baseline would be cheapest.

And on the pause arenas it was worse. The arena describes its pause tool as "…Call this and stop; you will be told the decision." All six frameworks were sending the first sentence only — dropping the instruction to pause, on the arena built to measure whether it pauses. Only vanilla kept it. Mock mode is blind for a sharper reason here: the script decides when the pause happens, so human_in_the_loop read 12/12 for six adapters while five were being told materially less than the sixth. → tool-schemas.md

The docs went on publishing a number after it had been withdrawn. The tool-schema correction above rewrote this page, overhead.md, the decision guide and the README — and missed four per-framework pages. Two of them (pydantic-ai.md at 0.96×, microsoft-agent-framework.md at 0.97×) kept asserting that a framework is cheaper than the hand-rolled baseline, the exact claim that had just been retracted, for five iterations. Nothing caught it because tests/test_doc_links.py gates that a cross-reference resolves, and nothing gated that a number is still true — a worse failure, because a stale figure reads as current.

tests/test_published_numbers.py now holds every page that quotes a tool_use overhead figure to the run record, and pins which pages must carry one so deleting or rewording a claim fails rather than quietly reducing coverage. Superseded numbers stay on the page inside a > block, which the gate skips deliberately: a benchmark that cannot write down that a number changed will stop writing it down. Half of it needs no framework installed — that the pages agree with each other — which is the half that would have caught this one. → tests/test_published_numbers.py

One of the probes on this page was contaminated by its own corpus. The batched-call measurement in §2b originally counted the word "observation" in the messages a framework returned — and a corpus entry describes Tokyo Tower as a "communications and observation tower", so it read 3 for a batch of 2. The published direction survived, but the wording did not: "the model is never told" was wrong for smolagents, which does tell the model and instead discards the successful sibling. The probe now matches each outcome's own text, and a test pins that none of those markers appear in the corpus — so a future corpus entry mentioning an error cannot quietly reintroduce the same drift.

Three frameworks' loop caps were not loop caps. max_tool_iterations is one shared knob, and measured against a mock that never stops asking for tools, a budget of 6 produced 50 calls (pydantic_ai, where retries is a tool-validation budget), 41 (microsoft_af, uncapped by default), and one call too many (smolagents). LangGraph needs 2 × recursion_limit. An adapter allowed to grind eight times longer posts better pass rates and incomparable latency, token and cost numbers — all four of which the scorecard publishes. Google ADK's RunConfig(max_llm_calls=N) is the only one that meant N out of the box.

Five adapters were ignoring the shared request timeout. request_timeout_s has existed since the first commit and only vanilla and langgraph ever passed it to their client. The other five inherited their library's default — ten minutes, on the official OpenAI client — and answered happily through a 20-second hang on a one-second budget. Same shape as the iteration-budget bug: a latency column would have been reporting library defaults rather than the arena's configuration. Measurable only because the mock learned to hang rather than fail. → transport.md

Temperature was an invariant with nothing behind it. methodology.md §1 says "temperature is 0.0 everywhere" in the same breath as one model for everyone — but it was ten hard-coded temperature=0.0 literals, one per adapter, and no check that any reached the gateway. Measured: all fourteen buildable adapters send 0.0 today, and no adapter sends any other sampling parameter — not top_p, seed, max_tokens, the penalties — so the provider default applies identically for everyone. smolagents alone adds stop sequences, which its text ReAct loop needs to parse its output. ArenaConfig.temperature is now the single source and tests/test_shared_controls.py holds every adapter to it on the wire with a non-zero canary, and holds every request body to the baseline's sampling envelope. A negative result, pinned the same way the model-on-wire check is. → fairness-controls.md

The mock's streaming path would have zeroed every token count. A real OpenAI-compatible provider sends no usage on a streamed response unless the client asks with stream_options: {"include_usage": true}. The mock sent none either way, so the first adapter to turn streaming on would have reported zero prompt and completion tokens with every correctness check still green — the answer is unaffected. Latent rather than live: no adapter here streams, which is itself worth recording, since the usual assumption is that agent frameworks stream by default and none of these does unless asked. The mock now honours the option, and the contract test permits streaming only alongside it.

→ methodology.md §3b

python -m pytest tests/test_usage_accounting.py tests/test_adapters_contract.py tests/test_mockserver.py -q

6. Things that turned out not to discriminate

Negative results are kept because they cost the same to produce and are the reason the positive ones are believable.

  • Batched tool calls, when all the calls are valid. All seven adapters run both and return both results. An arena was deliberately not built for this: an arena everybody passes adds runtime and dilutes the scorecard without separating anything. It lives as a test instead.
  • Graph orchestration, as a cost. See §3 — 1.00×. LangGraph's machinery adds nothing; vanilla_multi and langgraph_multi send the identical request. (It read 0.97× until a LangChain release equalised a tool-schema difference LangGraph used to carry.)
  • "Batching changes nothing." This was a hypothesis, and its test failed on langgraph — batching is not neutral. It was replaced with the invariant that does hold: batching may rescue a fault, but must never make a framework worse than it is serially.
  • Structured output — and this one is a finding about the arena. No adapter puts response_format on the wire; all seven ask for JSON in the prompt and hand back whatever arrives. Given a record that is not JSON, has wrong types, is missing a required field, or carries an extra one, every framework returns it byte-for-byte in the same two LLM calls a valid record costs. Nobody validates, nobody re-prompts, nobody repairs.

So structured_output is currently a tool_use arena whose answer is JSON-shaped, and arena/scorer.py is the only thing standing between a malformed record and a green scorecard — now gated, because nothing was pinning it. Stated rather than quietly fixed: giving one adapter a native mechanism and not the others would make the arena incomparable, so structured-output.md names what doing it properly requires instead.

python .github/scripts/report_structured_output.py
  • Retrieval, in mock mode. Every framework walks the rag arena identically — 15/15, exactly two hops on each multi-hop item, a passing refusal on each unanswerable one — because the script decides when the second search happens and what the final answer is. No adapter retrieves better here. What a framework could vary is the prompt cost of carrying a multi-document context, and that is the same per-request overhead already measured on tool_use in overhead.md. The arena's teeth are its dataset design (a multi-hop item a single lookup cannot answer) and its scorer (a hallucinated answer must fail an unanswerable item); both are gated in tests/test_rag_arena.py, and the cross-framework walk is now pinned there too as a floor.

What is not measured

Named explicitly, because the gap between "not measured" and "measured to be fine" is where benchmarks mislead:

  • Answer quality. Everything above is mock mode, where turns are scripted and identical for every framework. No finding here says any framework produces a better answer. That needs a live run and is blocked on an API key.
  • Latency and real cost. Same reason — mock-mode timings measure the harness.
  • Gemini through ADK, and the .NET half of Agent Framework. Out of scope: holding the model constant is what makes the comparison fair.
  • CrewAI, whose adapter is written but not yet mock-verified (needs Python 3.12), and Claude Agent SDK, still a stub.
  • Whether the delegation laws in §3 survive a real model. They are structural, so they should, but a real model may delegate more than once, or answer without delegating at all. Mock mode cannot tell you.
  • ~~Whether forwarding context between roles changes the ranking.~~ Measured — it does not; it widens the gap. See §3.
  • Whether any framework can be configured to manage context. §1b measures the default path, where nobody truncates. Several of these libraries expose hooks that would let you trim or summarise; none of them is on by default, and the arena does not turn any of them on.