Skip to content

Framework overhead

What this measures

Every adapter is handed the same arena prompt and the same two tool definitions, and the mock LLM replays byte-identical turns. So the size of the request each framework puts on the wire is not a property of the task or the model — it is the framework's own serialisation, and a provider bills for every byte of it.

This is one of only two things mock mode can compare honestly. (The other is resilience; see methodology.md §5.) Pass rate in mock mode is ~100% by construction and means nothing.

The numbers

Regenerated by CI on every run — the comparison job prints the live table. A local run on tool_use, 15 items, mean per item:

framework prompt tokens vs baseline
vanilla (stdlib baseline) 753.5 1.00×
langgraph 753.5 1.00×
pydantic_ai 794.0 1.05×
microsoft_af 802.0 1.06×
google_adk 836.1 1.11×
openai_agents 856.9 1.14×
smolagents 2935.5 3.90×

Completion tokens and LLM calls are identical across the first six (45.1 and 2.07), which is the point: the mock scripts the same turns for everyone, so the only thing that varies is the request. smolagents matches on LLM calls (2.07) but not on completion tokens (74.7) — see below.

These numbers changed once, for a reason worth knowing. An earlier version of this table had langgraph at 0.91×, and this page said prominently that the hand-rolled baseline was not the leanest. It was — the frameworks that measured cheaper were sending less: two dropped a parameter the arena declared, four dropped every parameter description. tool-schemas.md has the audit and the correction. With the tools genuinely equalised, langgraph ties the baseline byte for byte.

What drives it

Six of the seven are within a 1.15× band of each other, and smolagents sits right outside it. They are two different stories, so take them separately.

The 1.15× band

The conversation is identical — those six send exactly the same 472 characters of messages on the first turn, which is what PR #3's wire-level contract tests enforce. The whole spread comes from the tool block: the same two tools, described in the same OpenAI schema format and now carrying the same information, come out at

framework tools block what the extra bytes are
vanilla 637 chars the canonical spec, nothing added
langgraph 637 chars byte-identical to the baseline
pydantic_ai 715 chars additionalProperties: false
google_adk 735 chars title on every property, Args: folded into the description
microsoft_af 740 chars title on every property and the object
openai_agents 837 chars title, additionalProperties, and strict: true

Frameworks also differ in which control fields they send (openai_agents omits stream, langgraph omits tool_choice), but those are not billed and are excluded from the count.

Worth stating plainly, now that it is measured on equal terms: no framework is leaner on the wire than the hand-rolled loop. The baseline is the floor, and what separates the others from it is decoration — titles, additionalProperties, strict mode — rather than tighter serialisation. langgraph adds none of it and lands exactly on the floor.

That is a duller claim than the one it replaces, and it is the one the measurement supports.

smolagents: 3.90×, and it is all system prompt

smolagents is the exception that shows what the band is actually measuring. Its extra cost is not the tool schema — it is a templated system prompt that the framework prepends to whatever the arena asked for. The arena's own instruction is still in there verbatim (the contract tests check that), but it is under 10% of the 4,207 characters that go out, against vanilla's 384 — the arena's prompt and nothing else. That is ~3,800 extra characters resent on every request, and at 2.07 requests per item it accounts for essentially the whole gap.

Two details are worth pulling out:

  • The tools are transmitted twice. Every request carries both the 919-char OpenAI tools schema and a 1,159-char prose restatement of the same three tools inside the system prompt. Nothing reconciles them; a provider bills for both. This is also the reason smolagents moved up when the schemas were equalised: it pays for a restored parameter description in two places at once, which is why its multiple went from 3.77× to 3.90× while everyone else's moved by less.
  • final_answer is a real cost. smolagents ends its loop by calling a tool rather than by replying, so it must advertise a third tool the arena never declared. That is why its tools block is 919 chars against vanilla's 637, and it is why its completion tokens are 74.7 against everyone else's 45.1 — the final answer goes out as a JSON tool call instead of as plain content. Methodology §3 exempts control tools from the fairness rule; it does not exempt them from the bill.

None of this is a bug, and the prompt is doing real work — it is what lets smolagents drive weaker models through a tool loop without native tool-calling support. But if you are pairing it with a model that already tool-calls well, you are paying ~3.5 KB per request for scaffolding you do not need.

And this is the cheaper of its two agents. smolagents also ships CodeAgent, where the model writes and executes Python instead of emitting native tool calls. It runs as its own entry, smolagents_code, and it is heavier still — 6.95× the baseline prompt, because its system prompt is a longer few-shot (several worked Python examples) layered on top of everything above. Completion tokens go the other way (60.7 against ToolCallingAgent's 74.7 — a <code> blob is terser than a JSON tool call), and the LLM-call count is the same 2.07. The side-by-side table is on the smolagents page.

How to read these numbers

  • Relative, not absolute. The estimator is len(text) // 4, not a real BPE tokenizer, and JSON punctuation inflates it. Use the table to compare frameworks with each other; do not use it to predict a bill.
  • First-turn effects dominate on short tasks. The tool block is sent on every request, so its cost scales with the number of LLM calls, not with the length of the conversation. On a 2-call task it is a large fraction of the prompt; on a long one it matters much less — measured below.
  • This is per-arena. tool_use declares two tools; rag and structured_output declare one, so their spread is smaller. Run .github/scripts/report_overhead.py <arena> against any run record.

What happens when the loop gets long

The table above is a two-call task, which is where framework overhead looks its worst: a fixed per-request cost is being divided by two requests. The obvious next question is whether it ever stops mattering, and until now this page asserted an answer without measuring one.

Scripting the same conversation out to 30 tool-calling turns and recording every request:

request # 1 11 31
vanilla (stdlib baseline) 121 1531 4350
smolagents 1069 2551 5515
ratio 8.83× 1.67× 1.27×

(estimated prompt tokens; a smaller arena prompt than the table above, so compare the ratios and not the absolute numbers)

Regenerate with:

python .github/scripts/report_growth.py 30

Nobody truncates

Not one of the seven frameworks drops, windows or summarises anything. All 31 requests carry the entire history, and the curve is a straight line out to the last one. No adapter opted out of a feature here — none of these libraries ships context management on the default path.

So in a long tool loop, the thing bounding your prompt is max_tool_iterations, which is the harness's knob, not the library's. If you are running a real agent on a real bill, that is your problem to solve and no framework here solves it for you.

Frameworks differ by a constant, not by a rate

The marginal cost of one more turn, across all seven: 136.7 to 148.2 tokens. An 8% band, against an 8.8× spread on the first request. Adding a turn costs everyone the same; the frameworks differ only in where they start.

That is the whole mechanism behind the decay in the table. A per-request constant is paid once per request, so its cumulative cost grows linearly in turns, while the conversation it is being divided by grows quadratically — every turn resends every previous turn. Any fixed overhead therefore decays as 1/n.

Concretely, vanilla spends 9,901 estimated prompt tokens on a 10-turn item and 71,745 on a 30-turn one: 3× the turns, 7× the bill.

What that means for the headline number

smolagents' 3.90× is real, and it is a two-call number. Read it as an upper bound rather than a tax:

  • on short, high-volume items — classification, extraction, one-shot RAG — the multiple is close to the full 3.90× and worth caring about;
  • on a long agentic loop it converges toward 1×, and by turn 30 the framework you picked is a rounding error next to the number of turns you took.

The two properties above are gated in tests/test_prompt_growth.py — not the constants, which are findings, but the shape. If a framework started dropping history, every comparison on this page would be between different conversations, and a loud failure is better than a quietly wrong table.

History

Until the fix in arena/llm/mockserver.py, the mock estimated prompt_tokens from the messages array only. That understated every mock-mode prompt by roughly 1.8× — and because the tools block was the only thing that differed between frameworks, it reported all five as identical to within 1%. The comparison above did not exist; it was being silently discarded. tests/test_mockserver.py::test_prompt_tokens_include_the_tool_schemas guards the regression.