Findings¶
Everything this benchmark has actually measured, in one place, with the command that regenerates each number.
The per-framework pages and the topic docs are where the reasoning lives; this page is the index of results. If a claim is not on this page, it is not something the repo has measured.
How to read it¶
Three rules govern what appears here, and they are the whole reason the numbers are worth anything:
- Everything that could vary is held identical. Same model, same gateway, same tools, same task spec, same eval set, same iteration budget, and — in mock mode — byte-identical scripted turns. See methodology.md.
- Only two things compare honestly offline, and every finding below is one of them or a direct wire measurement: what a framework puts on the wire, and how it behaves when the model misbehaves. Mock-mode pass rate is ~100% by construction and means nothing. (methodology §5)
- Differences between frameworks are findings, not test failures. They are reported, not gated. What CI gates are the invariants — that adapters report their costs honestly, respect the budget, and pass the tool result through unaltered.
Every number below was regenerated by CI on a clean Linux install, not only locally.
1. What a framework costs you on the wire¶
Six of seven frameworks sit inside a 1.15× band; one sits at 3.90×.
Mean prompt tokens per item on tool_use, 15 items:
| framework | prompt tokens | vs baseline |
|---|---|---|
vanilla (stdlib baseline) |
753.5 | 1.00× |
| langgraph | 753.5 | 1.00× |
| pydantic_ai | 794.0 | 1.05× |
| microsoft_af | 802.0 | 1.06× |
| google_adk | 836.1 | 1.11× |
| openai_agents | 856.9 | 1.14× |
| smolagents | 2935.5 | 3.90× |
No framework is leaner on the wire than the hand-rolled loop — and this page
said the opposite until it was checked properly. See §5: the frameworks that
measured cheaper were sending less, and with the tool schemas equalised
langgraph ties the baseline byte for byte. Within the band the entire spread is
still the tools block — the messages are byte-identical at 472 characters —
but what separates the others from the floor is decoration (title,
additionalProperties, strict: true), not tighter serialisation.
smolagents' 3.90× is a system prompt, not tool serialisation. A 4,207-char template goes out on every request where the arena asked for 384, and 1,159 of those characters are a prose restatement of the same tools already sent as an OpenAI schema — the tools are transmitted twice and billed twice. It is doing real work (driving models that cannot natively tool-call), but paired with a model that tool-calls well it is ~3.8 KB per request of scaffolding you do not need.
Its CodeAgent shape is heavier again — 6.95×. smolagents_code runs the
same tools through a Python-writing loop instead of native tool calls; a longer
few-shot system prompt puts it well above the ToolCallingAgent entry, though
its completions run terser (60.7 tokens against 74.7). Same 2.07 LLM calls. The
side-by-side table is on
the smolagents page.
→ overhead.md · tool-schemas.md · smolagents.md
python -m arena run --arena tool_use --framework all --mode mock --no-scorecard && python .github/scripts/report_overhead.py tool_use
1b. …and what happens to that cost when the loop gets long¶
The table above is a two-call task, which is where framework overhead looks its worst. Running the same conversation out to 30 tool-calling turns:
| request # | 1 | 11 | 31 |
|---|---|---|---|
vanilla |
121 | 1531 | 4350 |
smolagents |
1069 | 2551 | 5515 |
| ratio | 8.83× | 1.67× | 1.27× |
Nobody truncates. Not one of the seven drops, windows or summarises anything;
all 31 requests carry the entire history. None of these libraries ships context
management on the default path, so the only thing bounding your prompt in a long
loop is max_tool_iterations — the harness's knob, not the library's.
Frameworks differ by a constant, not by a rate. One more turn costs everyone 136.7–148.2 tokens, an 8% band, against an 8.8× spread on the first request.
Those two facts explain the decay: a per-request constant is linear in turns,
the conversation it is divided by is quadratic, so any fixed overhead decays as
1/n. vanilla spends 9,901 estimated prompt tokens on a 10-turn item and
71,745 on a 30-turn one — 3× the turns, 7× the bill. Read 3.90× as an upper bound
on short items, not a tax on every item.
python .github/scripts/report_growth.py 30
2. What breaks when the model misbehaves¶
Eight scripted faults, byte-identical for every framework, so any difference is the framework's own error handling. This is the other thing mock mode compares honestly.
| framework | recovered | fails on / at what cost |
|---|---|---|
vanilla |
8/8 | — (2 LLM calls each) |
pydantic_ai |
8/8 | — |
microsoft_af |
8/8 | — |
langgraph |
7/8 | res-01 — malformed tool arguments |
openai_agents |
7/8 | res-02 — raises on an unknown tool name |
google_adk |
6/8 | res-01 and res-02, both uncaught exceptions |
smolagents |
8/8 | but 6 LLM calls (the whole budget) on the four faults its validator rejects before the tool runs, against 2 elsewhere |
smolagents' rough edge is structural, and now shows up as cost rather than a
lost item. The split is exact: on the four faults where the tool ran and
returned something it recovers in two calls like everyone else; on the four its
validation layer rejects first (unknown name, missing argument, unexpected
argument, null arguments) the error never reaches the transcript, the model
cannot see it, and it re-emits the identical call until the step budget is gone —
six calls and ~2.7× the prompt tokens, with zero tool calls recorded. It
still produces the right final answer at the end, so the item passes; a retry
wrapper does not help, because the prompt is byte-identical each time.
Correction. This row read 4/8 for several iterations. That was measured when an exhausted
smolagentsrun returned""; it now surfaces a final answer from its last memory step instead, sonumeric_equalspasses. The failure mechanism is unchanged — the validator-rejected faults still burn the whole budget with no tool call — but the consequence is a 3× cost, not a lost item.check_resilience.pynow gates every count in this table so the next drift fails CI instead of ageing silently.
Google ADK is the only framework that loses both res-01 and res-02, and
both are uncaught exceptions rather than the model giving up. Loud, at least, and
the kind of thing a retry wrapper does handle.
→ decision-guide.md §2 · feature-matrix.md
python -m arena run --arena resilience --framework all --mode mock --no-scorecard && python .github/scripts/check_resilience.py
2b. A quieter failure: what happens to a batch containing one broken call¶
When a model returns several tool calls in one turn and one of them is broken, two questions of the next request — did the successful call's result reach the model, and was the broken call reported at all?
| unknown tool | malformed args | missing required arg | |
|---|---|---|---|
vanilla, pydantic_ai, microsoft_af |
both | both | both |
langgraph |
both | good only | both |
smolagents |
error only | good only | error only |
openai_agents |
raises | both | both |
google_adk |
raises | raises | both |
Three ways to mishandle a batch, and they are not equally bad. Silent partial — the broken call vanishes with no message of any kind, and the run answers from partial evidence without the model ever being told. The successful sibling discarded — smolagents surfaces the error, but as a rewritten task rather than a turn in the transcript, dropping the good result with the history, so the model re-emits the identical batch. Raises — loud, and the best of the three, because nothing is silently wrong.
This refines the res-01 result above. LangGraph losing malformed arguments
is conditional: alone, the call produces no tool message and the graph halts,
so the item fails visibly; batched with a call that succeeds, the graph carries
on and the drop becomes invisible. The visible failure is the better outcome.
python -m pytest tests/test_parallel_tool_calls.py -q
2c. When the gateway fails, not the model¶
Every arena scripts the model. None scripts the provider returning 429, 500 or 400 — which is what real deployments hit. Attempts that reached the wire:
| framework | 429 once | 429 ×3 | 500 once | 400 |
|---|---|---|---|---|
vanilla (stdlib baseline) |
raises (1) | raises (1) | raises (1) | raises (1) |
langgraph |
ok (2) | raises (2) | ok (2) | raises (1) |
pydantic_ai, openai_agents, microsoft_af, google_adk |
ok (2) | raises (3) | ok (2) | raises (1) |
smolagents |
ok (2) | ok (4, +2–4 min) | ok (2) | raises (1) |
This is the clearest answer in the repo to "what does a framework buy you?"
The hand-rolled loop has no retry at all: one 429 and the item is lost. Every
framework survives it. It is the first dimension where the baseline loses
outright — on prompt size it is mid-pack, on resilience joint best, on
delegation exactly as expensive as a graph library.
smolagents is the only one that survives three consecutive 429s — by sleeping for minutes on one item (measured five times at 139–225 s, so heavily jittered). Nothing in the scorecard shows it: the item passes. In a batch that is throughput quietly collapsing. Everyone else fails fast in ~1.3 s and hands you the error. Both are defensible; neither is documented.
LangGraph retries once where the others retry twice, and everyone correctly refuses to retry a 400.
Every framework that retries honours Retry-After exactly — 3.01 s for a
header of 3, 9.02 s for 9. But smolagents' long sleep is not shortened by it:
there are two retry layers, and only the inner one (the OpenAI client) listens.
The outer one is smolagents.models.Retrying, whose first backoff is
RETRY_WAIT(60) × 2 × (1 + random()) — 120 to 240 s, bracketing every
measurement, and it never consults the header.
But that minutes-long sleep is rate-limits only. Against a provider that
simply hangs, smolagents fails in 8.5 s like everyone else — its outer layer is
built with retry_predicate=is_rate_limit_error, a substring check for 429 /
rate limit in the exception text, and a timeout matches none of it.
A hung provider also caught a bug in this harness, not in the frameworks.
ArenaConfig.request_timeout_s has existed since the first commit, and five of
the seven adapters never passed it to their client — pydantic_ai,
openai_agents, microsoft_af, smolagents and google_adk each inherited
their library's default (ten minutes, on the official OpenAI client) and answered
happily through a 20-second hang on a one-second budget. Fixed, and gated twice:
that a hang is abandoned, and that doubling the budget doubles the wait, which
is what stops any hard-coded value passing.
The knob turns out to bound an attempt, not an item. A framework that retries
twice spends timeout × attempts + backoff, so the default 60 s is a
three-minute worst case per item.
python .github/scripts/report_transport.py
3. What delegation costs¶
multi_agent carries real three-role pipelines (researcher → writer → editor)
alongside the single-agent entries, with identical role wording, so three
questions come apart:
| comparison | prompt | LLM calls |
|---|---|---|
vanilla → vanilla_multi — three stages, framework-free |
2.50× | 2.00× |
langgraph → langgraph_multi — the same, inside a framework |
2.50× | 2.00× |
vanilla_multi → langgraph_multi — the graph machinery alone |
1.00× | 1.00× |
openai_agents → openai_agents_multi — native handoffs |
2.64× | 2.00× |
smolagents → smolagents_multi — sub-agent invoked as a tool |
3.93× | 3.00× |
pydantic_ai → pydantic_ai_multi — the same, hand-built |
3.57× | 3.00× |
The cost of multi-agent is the structure, not the framework. Three roles
double the LLM calls and 2.5× the prompt tokens, and cost that whether you build
them with a graph library or a for loop.
The graph machinery adds nothing — vanilla_multi and langgraph_multi
send the byte-identical request. langgraph used to carry a 48.6-token
tool-schema difference from vanilla (a 0.97× that read like "the framework is
cheaper"); a LangChain release equalised the serialisation, so it now matches
byte for byte here as well as on tool_use.
Correction. The row above read
langgraph_multi2.62× and the graph machinery 0.97×, andopenai_agents_multi2.76× /smolagents_multi4.03×, for several iterations. A LangChain /openai-agents/smolagentsrelease bump moved each one's serialisation;vanillaandpydantic_aiare unchanged, every call multiplier is unchanged, andreport_delegation.pynow gates all five ratios.
You pay for a handoff by advertising it, not by taking it. Model-decided delegation costs ~10% more prompt than the structural pipeline at the same call count. Decomposing that gap across one item's four requests:
| messages | tool schemas | |
|---|---|---|
vanilla_multi |
6767 | 722 |
openai_agents_multi |
6818 | 1495 |
| difference | +51 | +773 |
The transfer call and its result add 51 characters to the entire conversation.
The transfer_to_* schemas add 773 — 94% of the difference — because they ride
on every request whether or not anyone delegates. A supervisor offering N
handoffs pays for N schemas on every request in the run: the cost scales with the
options you offer, not the ones you use.
Three roles cost 2× the calls in every mechanism except one. A sub-agent invoked as a tool costs 3×, because its reply is a tool result rather than the end of the run: the manager is still running and has to produce its own final answer afterwards. Visible on the wire as six requests, down the chain and back up it — the last two are pure mechanism. So this cost grows with the depth of a chain rather than with the work in it. (The 3.93× is a floor: the mock forwards only the task, so the writer never receives the researcher's findings. With a speaker swap the transcript comes along free; with a sub-agent you pass context by hand and pay for it again.)
And it generalises to a second, structurally different mechanism — at 3.3× the
price. smolagents' managed_agents advertises a sub-agent as an ordinary tool
rather than as a transfer that swaps the speaker. Swept the same way as the
handoff, both curves are linear in delegates offered: managed_agents costs
~875 characters per delegate on every request (after a ~385-char one-off
preamble), the OpenAI Agents SDK's handoffs ~251, from the first one, with
no preamble and nothing added to the system prompt. The gap is the same one
behind smolagents' 3.90× in §1 — each managed_agent is described twice, as
a JSON schema and again in prose, while a transfer_to_<name> schema is a
name-templated stub with an empty parameter object and no description of the
target at all. Both frameworks bill you for options rather than actions; they do
not bill the same amount, or in the same place.
Measured from one role to five, across five implementations in four libraries, every one follows an exact law:
| roles | 1 | 2 | 3 | 4 | 5 | ||
|---|---|---|---|---|---|---|---|
handoffs (openai_agents) |
speaker swap | 2 | 3 | 4 | 5 | 6 | N + 1 |
sub_agents (google_adk) |
transfer, returns to parent | 2 | 4 | 5 | 6 | 7 | N + 2 |
managed_agents (smolagents) |
sub-agent as a tool | 2 | 4 | 6 | 8 | 10 | 2N |
AgentTool (google_adk) |
sub-agent as a tool | 2 | 4 | 6 | 8 | 10 | 2N |
| agent delegation (pydantic_ai) | sub-agent as a tool | 2 | 4 | 6 | 8 | 10 | 2N |
2N is the mechanism, not the library — three implementations in three libraries that share no code, agreeing at every depth. The third one settles it: Pydantic AI has no delegation feature, so its chain is an ordinary async tool whose body awaits a nested run, and it costs exactly 2N anyway. The price is in the shape — a nested run returns a value instead of ending the conversation — and no library choice avoids it. "Handoff" is not one thing — both the OpenAI SDK and ADK call theirs a transfer, and ADK's costs one call more because it returns control to the parent, which then speaks again.
And you pay in calls or in prompt, not both — the call law alone points the
wrong way. Inside ADK at four roles: sub_agents costs 6 calls and 7375 prompt
tokens; AgentTool costs 8 calls and 1722. A transfer keeps one conversation
every agent sees all of, so the prompt compounds (15.4× against 3× the calls); a
sub-agent starts fresh, so the prompt stays flat and the calls compound instead
(3.59× against 4×). Cheaper in calls is dearer in prompt, and prompt is usually
the larger bill.
Restarting only helps if there is little to re-send — and the third sub-agent-as-tool implementation is what proves that is the library rather than the mechanism. smolagents restarts each sub-agent fresh and still compounds (5.68×); Pydantic AI restarts them exactly the same way and lands on ADK's numbers instead — 1581 prompt tokens at four roles against ADK's 1722 and smolagents' 13441. So smolagents' outlier is its own ~4 KB templated system prompt, re-sent by every fresh sub-agent (the scaffolding from §1), and nothing about starting fresh. Note too which implementation is cheapest in prompt at four roles: the one that costs the most calls.
Forwarding context does not change the ranking — it widens the gap. The obvious objection to everything above is that every pipeline here is measured with the delegate told only the task, which favours whichever mechanisms carry the transcript for free. Handing the same payload to every framework, extra prompt tokens per item against forwarding nothing:
| forwarded characters | 0 | 277 | 553 | 1105 |
|---|---|---|---|---|
vanilla_multi, langgraph_multi — structural |
0 | 0 | 0 | 0 |
openai_agents_multi — speaker swap |
0 | 0 | 0 | 0 |
smolagents_multi — sub-agent as tool |
0 | 509 | 978 | 1916 |
pydantic_ai_multi — sub-agent as tool |
0 | 508 | 977 | 1915 |
Three of the four mechanisms forward for free — a structural pipeline shares one
transcript, and a speaker swap's transfer_to_<agent> tool declares no parameters
at all, so there is nowhere to put a payload. Only a sub-agent starting a fresh
conversation has to be told, and the two implementations of that agree to within
one token at every size, in libraries that share no code.
It costs about 1.77 tokens per forwarded character, not the 0.25 of a single
copy: the payload becomes the sub-agent's opening message and rides every
subsequent request of its conversation. So 3.93× and 3.57× were floors, and a
modest 553-character findings block takes pydantic_ai_multi to roughly 4.9×
while openai_agents_multi stays at 2.64×. The gap about doubles; it does not
invert.
Two things this deliberately does not say. The lower completion tokens on the handoff chain (140 vs 198) are not efficiency — the researcher emits a short transfer instead of a draft, so only the last agent writes a brief. And the mock scripts identical turns for all four entries, so the benefit is held at zero by construction: this is the bill delegation puts on the table before it has done anything for you.
python -m arena run --arena multi_agent --framework all --mode mock --no-scorecard && python .github/scripts/report_delegation.py
4. Pausing, and surviving a crash¶
Six frameworks pause; no two of the six mechanisms look alike.
human_in_the_loop observes the pause in the harness rather than trusting the
agent's prose, and durable_state throws the runner away at the pause and
rebuilds it, so only a real checkpoint gets across.
| adapter | pause | crash | mechanism |
|---|---|---|---|
langgraph |
12/12 | 8/8 | interrupt() + MemorySaver / SqliteSaver, resumed with Command(resume=…) |
openai_agents |
12/12 | 8/8 | needs_approval=True → ToolApprovalItem; RunState.to_json() serialises the whole run |
pydantic_ai |
12/12 | 8/8 | tool raises CallDeferred → DeferredToolRequests; stateless replay as message_history |
google_adk |
12/12 🟡 | 8/8 | LongRunningFunctionTool; DatabaseSessionService on sqlite+aiosqlite |
microsoft_af |
12/12 | ❌ | approval_mode="always_require" + ToolApprovalMiddleware |
vanilla |
12/12 🟡 | 8/8 | emulated — transcript carried back in; durable, but not a checkpointer |
Google ADK reports the pause but does not enforce it. Left alone, ADK hands
the model the {"status": "pending"} marker and the model carries on — measured,
it went straight to book_room with no human decision, which is exactly what the
arena's no_tool_before_suspend check exists to catch. Breaking out of the event
stream is what makes the pause real. It is the only one of the six where ignoring
the signal lets the agent act, hence 🟡.
microsoft_af's durable gap was measured, not assumed. Its approval state
serialises cleanly (~400 chars of clean JSON), but the conversation lives in the
AgentSession's in-memory store, which returns raw strings after a JSON round
trip, and restoring the approval state into a rebuilt agent re-queues the
request instead of consuming the answer. The adapter therefore does not expose
resume on a durable arena at all — keeping one it cannot honour would post 0/8
and read as a broken framework rather than an absent feature.
The durable_state check is not vacuous. An adapter patched to restart from
scratch instead of resuming drops to 0/8: it still reaches the right answer,
by redoing both lookups, and the call_counts check catches the duplicated work.
What the pause costs, and why "fits in a message list" is the whole point.
resume_state is the only thing carried across the restart. On the two-lookup
durable_state item:
| adapter | resume_state |
on-disk store | what is in the state |
|---|---|---|---|
google_adk |
112 B | 57 KB sqlite | a session id |
langgraph |
401 B | 76 KB sqlite | a thread id + config |
vanilla |
2.5 KB | — | the whole transcript |
pydantic_ai |
6.1 KB | — | message history as JSON |
openai_agents |
17 KB | — | RunState.to_json() — the entire run object |
A checkpointer's resume_state is a handle — a few hundred bytes whatever
the conversation did — because the run lives in its store (a fixed sqlite floor,
not proportional to turns). A serialiser's resume_state is the
conversation, so it grows with every turn, and openai_agents is heaviest
because it serialises the run object rather than just the messages. Both survive
the restart; only the first kind still does once the transcript is large. The
under-1-KB handle and the several-KB transcript are gated as a split in
tests/test_durable_across_a_restart.py so a checkpointer that started inlining
the transcript would fail.
→ feature-matrix.md · per-framework pages in frameworks/
python -m arena run --arena human_in_the_loop --framework all --mode mock --no-scorecard
python -m arena run --arena durable_state --framework all --mode mock --no-scorecard
5. Can the numbers themselves be trusted?¶
Two findings that are about the benchmark rather than about any framework. Both came from asking what nothing was checking.
durable_state could be passed by cheating, and now cannot. The arena's
claim is that the runner is thrown away at the pause. The harness enforced that
by rebuilding the runner and JSON round-tripping the state — both inside one
interpreter, which a module-level cache keyed by item id survives. Measured
rather than supposed: a wrapper doing exactly that scores 8/8, with
['search', 'search', 'calculator'] and one suspend, indistinguishable in the
scorecard from the honest baseline it wraps.
Durability now runs its two legs in two processes, with nothing between them
but a JSON file and checkpoint_dir. All five resumable adapters pass — a
negative result, and the reason the published 8/8s mean what they say — while the
cheat raises KeyError. The cheat ships as a test fixture so a green file cannot
come to mean "the probe ran and asserted nothing".
It also settles the stateless-resume / real-checkpointer split by measurement
rather than by reading the source: langgraph writes langgraph.sqlite and
google_adk writes adk_sessions.sqlite; vanilla, pydantic_ai and
openai_agents write nothing and carry the whole conversation in
resume_state. → methodology.md §7
Nothing was verifying that adapters report cost honestly — every published
comparison on this page assumed it. A gate holding each AgentResult against
what the mock server actually served found a real bug immediately: langgraph
reported 3 LLM calls on durable_state where 4 were served, dropping a whole
leg. Correctness was never affected (8/8 before and after), which is precisely
why it survived. The same gate then caught a second, different accounting bug in
the very next adapter written against it. Every other adapter verified honest —
a negative result, and the thing that makes §1 and §3 worth reading.
The frameworks were not describing the same tool. §1 said for fifteen
iterations that the hand-rolled baseline is not the leanest — that three of
four frameworks serialise the same tools more compactly. Nothing had compared
what those bytes say, only what they cost. Against the canonical spec in
arena/tools/__init__.py:
google_adkandsmolagentsdeclaredsearch(query)where the arena declaressearch(query, k=3)— a parameter the model was never offered;langgraph,pydantic_ai,openai_agentsandmicrosoft_afsent bare types, no parameter descriptions, because none of these libraries reads one from the docstring and each wants it somewhere different.
Mock mode is structurally blind to this: scripted calls are replayed whatever the schema said, so every arena stayed green. Live it would have read as a quality difference and been blamed on the framework.
With the tools equalised, langgraph ties the baseline exactly — 753.5
tokens against 753.5, tools block 637 chars against 637. Its whole 0.91×
advantage was four missing descriptions, and the corrected claim is the duller
one: no framework is leaner than the hand-rolled loop, and the band is 1.00×–1.14×
rather than 0.91×–1.05×. The second correction to that claim, having already
replaced an earlier hypothesis that the baseline would be cheapest.
And on the pause arenas it was worse. The arena describes its pause tool as
"…Call this and stop; you will be told the decision." All six frameworks
were sending the first sentence only — dropping the instruction to pause, on
the arena built to measure whether it pauses. Only vanilla kept it. Mock mode
is blind for a sharper reason here: the script decides when the pause happens,
so human_in_the_loop read 12/12 for six adapters while five were being told
materially less than the sixth. → tool-schemas.md
The docs went on publishing a number after it had been withdrawn. The
tool-schema correction above rewrote this page, overhead.md, the decision guide
and the README — and missed four per-framework pages. Two of them
(pydantic-ai.md at 0.96×, microsoft-agent-framework.md at 0.97×) kept
asserting that a framework is cheaper than the hand-rolled baseline, the exact
claim that had just been retracted, for five iterations. Nothing caught it because
tests/test_doc_links.py gates that a cross-reference resolves, and nothing
gated that a number is still true — a worse failure, because a stale figure
reads as current.
tests/test_published_numbers.py now holds every page that quotes a tool_use
overhead figure to the run record, and pins which pages must carry one so
deleting or rewording a claim fails rather than quietly reducing coverage.
Superseded numbers stay on the page inside a > block, which the gate skips
deliberately: a benchmark that cannot write down that a number changed will stop
writing it down. Half of it needs no framework installed — that the pages agree
with each other — which is the half that would have caught this one.
→ tests/test_published_numbers.py
One of the probes on this page was contaminated by its own corpus. The
batched-call measurement in §2b originally counted the word "observation" in the
messages a framework returned — and a corpus entry describes Tokyo Tower as a
"communications and observation tower", so it read 3 for a batch of 2. The
published direction survived, but the wording did not: "the model is never told"
was wrong for smolagents, which does tell the model and instead discards the
successful sibling. The probe now matches each outcome's own text, and a test
pins that none of those markers appear in the corpus — so a future corpus entry
mentioning an error cannot quietly reintroduce the same drift.
Three frameworks' loop caps were not loop caps. max_tool_iterations is one
shared knob, and measured against a mock that never stops asking for tools, a
budget of 6 produced 50 calls (pydantic_ai, where retries is a
tool-validation budget), 41 (microsoft_af, uncapped by default), and one call
too many (smolagents). LangGraph needs 2 × recursion_limit. An adapter
allowed to grind eight times longer posts better pass rates and incomparable
latency, token and cost numbers — all four of which the scorecard publishes.
Google ADK's RunConfig(max_llm_calls=N) is the only one that meant N out of
the box.
Five adapters were ignoring the shared request timeout. request_timeout_s
has existed since the first commit and only vanilla and langgraph ever passed
it to their client. The other five inherited their library's default — ten
minutes, on the official OpenAI client — and answered happily through a
20-second hang on a one-second budget. Same shape as the iteration-budget bug: a
latency column would have been reporting library defaults rather than the arena's
configuration. Measurable only because the mock learned to hang rather than
fail. → transport.md
Temperature was an invariant with nothing behind it. methodology.md §1 says
"temperature is 0.0 everywhere" in the same breath as one model for everyone —
but it was ten hard-coded temperature=0.0 literals, one per adapter, and no
check that any reached the gateway. Measured: all fourteen buildable adapters send
0.0 today, and no adapter sends any other sampling parameter — not top_p,
seed, max_tokens, the penalties — so the provider default applies identically
for everyone. smolagents alone adds stop sequences, which its text ReAct loop
needs to parse its output. ArenaConfig.temperature is now the single source and
tests/test_shared_controls.py holds every adapter to it on the wire with a
non-zero canary, and holds every request body to the baseline's sampling envelope.
A negative result, pinned the same way the model-on-wire check is. →
fairness-controls.md
The mock's streaming path would have zeroed every token count. A real
OpenAI-compatible provider sends no usage on a streamed response unless the
client asks with stream_options: {"include_usage": true}. The mock sent none
either way, so the first adapter to turn streaming on would have reported zero
prompt and completion tokens with every correctness check still green — the
answer is unaffected. Latent rather than live: no adapter here streams, which
is itself worth recording, since the usual assumption is that agent frameworks
stream by default and none of these does unless asked. The mock now honours the
option, and the contract test permits streaming only alongside it.
python -m pytest tests/test_usage_accounting.py tests/test_adapters_contract.py tests/test_mockserver.py -q
6. Things that turned out not to discriminate¶
Negative results are kept because they cost the same to produce and are the reason the positive ones are believable.
- Batched tool calls, when all the calls are valid. All seven adapters run both and return both results. An arena was deliberately not built for this: an arena everybody passes adds runtime and dilutes the scorecard without separating anything. It lives as a test instead.
- Graph orchestration, as a cost. See §3 — 1.00×. LangGraph's machinery
adds nothing;
vanilla_multiandlanggraph_multisend the identical request. (It read 0.97× until a LangChain release equalised a tool-schema difference LangGraph used to carry.) - "Batching changes nothing." This was a hypothesis, and its test failed on
langgraph— batching is not neutral. It was replaced with the invariant that does hold: batching may rescue a fault, but must never make a framework worse than it is serially. - Structured output — and this one is a finding about the arena. No adapter
puts
response_formaton the wire; all seven ask for JSON in the prompt and hand back whatever arrives. Given a record that is not JSON, has wrong types, is missing a required field, or carries an extra one, every framework returns it byte-for-byte in the same two LLM calls a valid record costs. Nobody validates, nobody re-prompts, nobody repairs.
So structured_output is currently a tool_use arena whose answer is
JSON-shaped, and arena/scorer.py is the only thing standing between a
malformed record and a green scorecard — now gated, because nothing was
pinning it. Stated rather than quietly fixed: giving one adapter a native
mechanism and not the others would make the arena incomparable, so
structured-output.md names what doing it properly
requires instead.
python .github/scripts/report_structured_output.py
- Retrieval, in mock mode. Every framework walks the
ragarena identically — 15/15, exactly two hops on each multi-hop item, a passing refusal on each unanswerable one — because the script decides when the second search happens and what the final answer is. No adapter retrieves better here. What a framework could vary is the prompt cost of carrying a multi-document context, and that is the same per-request overhead already measured ontool_usein overhead.md. The arena's teeth are its dataset design (a multi-hop item a single lookup cannot answer) and its scorer (a hallucinated answer must fail an unanswerable item); both are gated intests/test_rag_arena.py, and the cross-framework walk is now pinned there too as a floor.
What is not measured¶
Named explicitly, because the gap between "not measured" and "measured to be fine" is where benchmarks mislead:
- Answer quality. Everything above is mock mode, where turns are scripted and identical for every framework. No finding here says any framework produces a better answer. That needs a live run and is blocked on an API key.
- Latency and real cost. Same reason — mock-mode timings measure the harness.
- Gemini through ADK, and the
.NEThalf of Agent Framework. Out of scope: holding the model constant is what makes the comparison fair. - CrewAI, whose adapter is written but not yet mock-verified (needs Python 3.12), and Claude Agent SDK, still a stub.
- Whether the delegation laws in §3 survive a real model. They are structural, so they should, but a real model may delegate more than once, or answer without delegating at all. Mock mode cannot tell you.
- ~~Whether forwarding context between roles changes the ranking.~~ Measured — it does not; it widens the gap. See §3.
- Whether any framework can be configured to manage context. §1b measures the default path, where nobody truncates. Several of these libraries expose hooks that would let you trim or summarise; none of them is on by default, and the arena does not turn any of them on.