Framework overhead¶
What this measures¶
Every adapter is handed the same arena prompt and the same two tool definitions, and the mock LLM replays byte-identical turns. So the size of the request each framework puts on the wire is not a property of the task or the model — it is the framework's own serialisation, and a provider bills for every byte of it.
This is one of only two things mock mode can compare honestly. (The other is
resilience; see methodology.md §5.) Pass rate in mock mode is
~100% by construction and means nothing.
The numbers¶
Regenerated by CI on every run — the comparison job prints the live table.
A local run on tool_use, 15 items, mean per item:
| framework | prompt tokens | vs baseline |
|---|---|---|
vanilla (stdlib baseline) |
753.5 | 1.00× |
| langgraph | 753.5 | 1.00× |
| pydantic_ai | 794.0 | 1.05× |
| microsoft_af | 802.0 | 1.06× |
| google_adk | 836.1 | 1.11× |
| openai_agents | 856.9 | 1.14× |
| smolagents | 2935.5 | 3.90× |
Completion tokens and LLM calls are identical across the first six (45.1 and
2.07), which is the point: the mock scripts the same turns for everyone, so the
only thing that varies is the request. smolagents matches on LLM calls (2.07)
but not on completion tokens (74.7) — see below.
These numbers changed once, for a reason worth knowing. An earlier version of this table had
langgraphat 0.91×, and this page said prominently that the hand-rolled baseline was not the leanest. It was — the frameworks that measured cheaper were sending less: two dropped a parameter the arena declared, four dropped every parameter description. tool-schemas.md has the audit and the correction. With the tools genuinely equalised,langgraphties the baseline byte for byte.
What drives it¶
Six of the seven are within a 1.15× band of each other, and smolagents sits
right outside it. They are two different stories, so take them separately.
The 1.15× band¶
The conversation is identical — those six send exactly the same 472 characters
of messages on the first turn, which is what PR #3's wire-level contract tests
enforce. The whole spread comes from the tool block: the same two tools,
described in the same OpenAI schema format and now carrying the same
information, come out at
| framework | tools block |
what the extra bytes are |
|---|---|---|
vanilla |
637 chars | the canonical spec, nothing added |
| langgraph | 637 chars | byte-identical to the baseline |
| pydantic_ai | 715 chars | additionalProperties: false |
| google_adk | 735 chars | title on every property, Args: folded into the description |
| microsoft_af | 740 chars | title on every property and the object |
| openai_agents | 837 chars | title, additionalProperties, and strict: true |
Frameworks also differ in which control fields they send (openai_agents omits
stream, langgraph omits tool_choice), but those are not billed and are
excluded from the count.
Worth stating plainly, now that it is measured on equal terms: no framework is
leaner on the wire than the hand-rolled loop. The baseline is the floor, and
what separates the others from it is decoration — titles, additionalProperties,
strict mode — rather than tighter serialisation. langgraph adds none of it and
lands exactly on the floor.
That is a duller claim than the one it replaces, and it is the one the measurement supports.
smolagents: 3.90×, and it is all system prompt¶
smolagents is the exception that shows what the band is actually measuring.
Its extra cost is not the tool schema — it is a templated system prompt that
the framework prepends to whatever the arena asked for. The arena's own
instruction is still in there verbatim (the contract tests check that), but it is
under 10% of the 4,207 characters that go out, against vanilla's 384 — the
arena's prompt and nothing else. That is ~3,800 extra characters resent on every
request, and at 2.07 requests per item it accounts for essentially the whole gap.
Two details are worth pulling out:
- The tools are transmitted twice. Every request carries both the 919-char
OpenAI
toolsschema and a 1,159-char prose restatement of the same three tools inside the system prompt. Nothing reconciles them; a provider bills for both. This is also the reasonsmolagentsmoved up when the schemas were equalised: it pays for a restored parameter description in two places at once, which is why its multiple went from 3.77× to 3.90× while everyone else's moved by less. final_answeris a real cost.smolagentsends its loop by calling a tool rather than by replying, so it must advertise a third tool the arena never declared. That is why itstoolsblock is 919 chars againstvanilla's 637, and it is why its completion tokens are 74.7 against everyone else's 45.1 — the final answer goes out as a JSON tool call instead of as plain content. Methodology §3 exempts control tools from the fairness rule; it does not exempt them from the bill.
None of this is a bug, and the prompt is doing real work — it is what lets
smolagents drive weaker models through a tool loop without native tool-calling
support. But if you are pairing it with a model that already tool-calls well, you
are paying ~3.5 KB per request for scaffolding you do not need.
And this is the cheaper of its two agents. smolagents also ships
CodeAgent, where the model writes and executes Python instead of emitting
native tool calls. It runs as its own entry, smolagents_code, and it is
heavier still — 6.95× the baseline prompt, because its system prompt is a
longer few-shot (several worked Python examples) layered on top of everything
above. Completion tokens go the other way (60.7 against ToolCallingAgent's
74.7 — a <code> blob is terser than a JSON tool call), and the LLM-call count
is the same 2.07. The side-by-side table is on
the smolagents page.
How to read these numbers¶
- Relative, not absolute. The estimator is
len(text) // 4, not a real BPE tokenizer, and JSON punctuation inflates it. Use the table to compare frameworks with each other; do not use it to predict a bill. - First-turn effects dominate on short tasks. The tool block is sent on every request, so its cost scales with the number of LLM calls, not with the length of the conversation. On a 2-call task it is a large fraction of the prompt; on a long one it matters much less — measured below.
- This is per-arena.
tool_usedeclares two tools;ragandstructured_outputdeclare one, so their spread is smaller. Run.github/scripts/report_overhead.py <arena>against any run record.
What happens when the loop gets long¶
The table above is a two-call task, which is where framework overhead looks its worst: a fixed per-request cost is being divided by two requests. The obvious next question is whether it ever stops mattering, and until now this page asserted an answer without measuring one.
Scripting the same conversation out to 30 tool-calling turns and recording every request:
| request # | 1 | 11 | 31 |
|---|---|---|---|
vanilla (stdlib baseline) |
121 | 1531 | 4350 |
smolagents |
1069 | 2551 | 5515 |
| ratio | 8.83× | 1.67× | 1.27× |
(estimated prompt tokens; a smaller arena prompt than the table above, so compare the ratios and not the absolute numbers)
Regenerate with:
python .github/scripts/report_growth.py 30
Nobody truncates¶
Not one of the seven frameworks drops, windows or summarises anything. All 31 requests carry the entire history, and the curve is a straight line out to the last one. No adapter opted out of a feature here — none of these libraries ships context management on the default path.
So in a long tool loop, the thing bounding your prompt is max_tool_iterations,
which is the harness's knob, not the library's. If you are running a real agent
on a real bill, that is your problem to solve and no framework here solves it
for you.
Frameworks differ by a constant, not by a rate¶
The marginal cost of one more turn, across all seven: 136.7 to 148.2 tokens. An 8% band, against an 8.8× spread on the first request. Adding a turn costs everyone the same; the frameworks differ only in where they start.
That is the whole mechanism behind the decay in the table. A per-request constant
is paid once per request, so its cumulative cost grows linearly in turns,
while the conversation it is being divided by grows quadratically — every
turn resends every previous turn. Any fixed overhead therefore decays as 1/n.
Concretely, vanilla spends 9,901 estimated prompt tokens on a 10-turn item and
71,745 on a 30-turn one: 3× the turns, 7× the bill.
What that means for the headline number¶
smolagents' 3.90× is real, and it is a two-call number. Read it as an upper
bound rather than a tax:
- on short, high-volume items — classification, extraction, one-shot RAG — the multiple is close to the full 3.90× and worth caring about;
- on a long agentic loop it converges toward 1×, and by turn 30 the framework you picked is a rounding error next to the number of turns you took.
The two properties above are gated in
tests/test_prompt_growth.py — not the
constants, which are findings, but the shape. If a framework started dropping
history, every comparison on this page would be between different conversations,
and a loud failure is better than a quietly wrong table.
History¶
Until the fix in arena/llm/mockserver.py, the mock estimated prompt_tokens
from the messages array only. That understated every mock-mode prompt by
roughly 1.8× — and because the tools block was the only thing that differed
between frameworks, it reported all five as identical to within 1%. The
comparison above did not exist; it was being silently discarded.
tests/test_mockserver.py::test_prompt_tokens_include_the_tool_schemas guards
the regression.