smolagents¶
Adapter · smolagents[openai]==1.26.0 ·
runs 5 of 7 arenas
Hugging Face's minimal agent library. The smolagents entry uses
ToolCallingAgent (the native-tool-calling agent), which is what the six-way
overhead and resilience comparisons rank. CodeAgent — the model writes and
executes Python — is a different task shape, and it has its own contrast entry,
smolagents_code (the three single-agent tool arenas — tool_use, rag,
structured_output); its cost is in the prompt-size section
below.
Wiring¶
model = OpenAIServerModel(
model_id=config.model,
api_base=config.base_url,
api_key=config.api_key,
temperature=0.0,
)
agent = ToolCallingAgent(
tools=_make_tools(_tool_names(arena.tools)),
model=model,
instructions=arena.system_prompt,
max_steps=max(1, config.max_tool_iterations - 1),
verbosity_level=0,
)
Four things cost real debugging time.
The [openai] extra is not optional¶
pip install smolagents gives you a package where OpenAIServerModel raises at
construction. The openai extra is what supplies it. Pinned in
requirements.txt with that note,
because the failure mode reads like a bad import, not a missing extra.
max_steps is off by one¶
smolagents makes one model call beyond max_steps — a final attempt once the
budget is spent. On a budget of 6 it sent 7 requests. max_steps = N - 1 yields
exactly the N LLM calls every other adapter gets, which the contract test in
tests/test_adapters_contract.py enforces against a mock that never stops asking
for tools.
Token usage: sum the steps, do not diff the monitor¶
agent.monitor looks like the obvious source, but run(reset=True) calls
monitor.reset(), so a before/after subtraction silently produces garbage (it
gave 2639 / 0 / 4 on a run that really cost far less). Summing each step's
step.token_usage is both correct and strictly better: it also counts steps whose
tool call failed, which is exactly the cost resilience is trying to measure.
An exhausted run returns "", not an exception¶
When the step budget runs out, agent.run() returns an empty string and the
error lives on the last memory step. Without surfacing it the arena records a
silent blank answer. The adapter promotes the last step.error to
AgentResult.error when the output is empty, which is how res-02 below reads as
AgentMaxStepsError: Reached max steps. instead of as nothing at all.
Two protocol differences the harness had to accommodate¶
Both are framework facts, not adapter choices, and both are handled the same way the CrewAI text-ReAct accommodation was — by teaching the mock to speak the client's protocol, never by giving the agent a capability the arena withheld.
It ends the loop by calling a tool. smolagents advertises its own
final_answer tool and reads a plain content reply as "not finished yet".
Un-accommodated, every item burned its whole step budget where 2 requests were
scripted. The mock now renders a scripted content turn as a final_answer call
for any client that advertises one
(arena.llm.mockserver._wants_final_answer_tool), and arena.tools.CONTROL_TOOLS
keeps that call out of the adapter's reported tool_calls so it cannot satisfy a
tool_used check or blow a max_tool_calls cap. See
methodology.md §3.
Tool results come back as user messages. Not role: "tool", and with list
content-parts rather than a string. The wire-level contract tests were keyed on
role == "tool"; they now locate the tool result by content and flatten
content-parts, so they still assert the same thing — that the tool's output
reaches the model unaltered — for every adapter.
Results (comparable columns marked)¶
| arena | result |
|---|---|
tool_use |
15/15 |
structured_output |
15/15 |
rag |
15/15 |
multi_agent |
10/10 |
resilience |
8/8, at 3× the cost on four of them (comparable) |
human_in_the_loop |
unsupported — no resume method |
durable_state |
unsupported — no resume method |
Mock pass rates are ~100% by construction and prove only correct wiring. What
compares frameworks honestly is prompt size and what a fault costs — the
resilience pass rate is now 8/8 for smolagents, so the signal moved into the
per-item LLM-call count.
Prompt size: 3.90× baseline (comparable)¶
The other six adapters sit within a 1.15× band. smolagents is at 3.90×
(2936 estimated prompt tokens per item against vanilla's 754), and the cause is
entirely its templated system prompt — 3932 characters where the arena asked for
384, resent on every request. The arena's instruction survives verbatim inside it,
but it is under 10% of what goes out; the rest is a framework preamble, two worked
examples, a trailing rules block, and a prose restatement of the same tools that
are already in the OpenAI tools schema. Full breakdown in
overhead.md.
Its completion tokens are also higher (74.7 vs 45.1) for a related reason: the final answer leaves as a JSON tool call rather than as plain content.
This is not waste in the abstract — that scaffolding is what drives weaker models through a tool loop without native tool-calling support. It is waste if you pair it with a model that already tool-calls well.
And CodeAgent is heavier still — 6.95×. frameworks/smolagents_code runs
the same library and tools as a CodeAgent: the model replies with a Python
<code> block that calls the tool functions, which smolagents executes.
| entry | prompt tok | vs base | completion |
|---|---|---|---|
smolagents (ToolCallingAgent) |
2936 | 3.90× | 74.7 |
smolagents_code (CodeAgent) |
5240 | 6.95× | 60.7 |
CodeAgent's system prompt is a longer few-shot — several worked examples of
writing Python against notional tools — so it puts ~1.8× the ToolCallingAgent
prompt on every request. Its completion tokens are lower (60.7 vs 74.7): a
<code> blob is terser than a JSON tool call plus reasoning. Same 2.07 LLM
calls, because the scripted turns are identical. If you reach for CodeAgent
for its execution model, this is the wire bill it comes with.
resilience: 8/8, but four of them cost 3× the rest¶
Every fault is eventually answered — but the eight divide perfectly on one question, does the failure reach the transcript?, and that line is the cost.
| item | fault | recorded? | LLM calls |
|---|---|---|---|
res-01 |
malformed JSON arguments | yes | 2 |
res-03 |
tool ran, returned ERROR |
yes | 2 |
res-06 |
tool ran, returned No results. |
yes | 2 |
res-07 |
tool ran, evaluator refused | yes | 2 |
res-02 |
tool name does not exist | no | 6 |
res-04 |
required argument missing | no | 6 |
res-05 |
argument not in the schema | no | 6 |
res-08 |
arguments serialised as null |
no | 6 |
The dividing line is smolagents' own tool-validation layer. Anything that
gets as far as running the tool comes back as an observation and lands in the
conversation; anything the dispatcher rejects first — unknown name, missing
argument, unexpected argument, null arguments — raises and is dropped.
The evidence is on the wire. On all four of those, every request after the first
sent exactly ['system', 'user'] — the prompt never changes, so the model emits
the identical bad call until the step budget is gone, verbatim:
Argument expr is required
Argument expr is required
Argument expr is required
Argument expr is required
Argument expr is required
Reached max steps.
smolagents then returns its last memory step, which happens to contain the
answer the mock served, so numeric_equals passes — that is why this row reads
8/8 and not 4/8. What it spent getting there is six LLM calls and ~2.7× the
prompt tokens against the two calls the other four faults take.
On the four recorded faults the second request carried
['system', 'user', 'assistant', 'user'] — the attempt was there, the model
could see what it had done, and it corrected on the next turn. res-01 is the
proof it is feedback and not severity: malformed JSON is the least structured
failure of the eight and it recovers cheaply, because it is caught at parse time
and written back.
Correction. This section read
4/8for several iterations. The mechanism is unchanged — the validator-rejected faults still burn the whole budget with no tool call — butsmolagentsnow surfaces a final answer from its last memory step rather than returning"", so the item passes at a cost instead of failing.check_resilience.pygates the count now.
managed_agents: the same design decision, again¶
managed_agents advertises a sub-agent as an ordinary tool named after itself;
calling it runs a whole nested agent and returns its answer. Measured with no
delegation happening at all, characters of system prompt plus tool schema:
| sub-agents offered | system prompt | tool schemas | total | marginal |
|---|---|---|---|---|
| 0 | 3311 | 518 | 3829 | — |
| 1 | 4102 | 982 | 5084 | +1255 |
| 2 | 4498 | 1455 | 5953 | +869 |
| 3 | 4900 | 1934 | 6834 | +881 |
~875 characters per offered sub-agent, on every request, whether or not anyone delegates. (The larger first step is a ~385-char preamble that switches delegation on and is paid once.)
The marginal cost splits into ~400 characters of prose and ~475 of JSON schema, because each sub-agent is described twice — the same double transmission as the tools above, in a second place. It is the single most consistent thing about this framework's wire behaviour.
For scale: the OpenAI Agents SDK charges 262 characters to offer a handoff, so this is 3.3× to hold open the same option.
Taking the delegation costs more too — and here that is the framework's
architecture, not its prompting. smolagents_multi runs the same three roles
as the other pipeline entries and spends 6 model calls where they spend 4
(3.93× / 3.00× against the single-agent entry, where every other mechanism is
~2.5× / 2.00×). A sub-agent's reply is a tool result, not the end of the run,
so the manager is still going and has to produce its own final answer once the
sub-agent returns. Every level of nesting costs one extra call. See
multi-agent.md.
Context does not travel either: a managed sub-agent starts a fresh conversation and receives only the task string, where a speaker-swapping handoff inherits the whole transcript. Passing findings down the chain is something you do by hand, and pay for again.
Two retry layers, and the outer one does not listen¶
OpenAIServerModel wraps the OpenAI client in smolagents.models.Retrying, so a
rate-limited call is retried at two levels:
| layer | attempts | first backoff | honours Retry-After? |
|---|---|---|---|
| the OpenAI client | 2 retries | ~0.5 s | yes |
Retrying |
RETRY_MAX_ATTEMPTS = 3 |
RETRY_WAIT(60) x 2 x (1 + random()) = 120-240 s |
no |
The effect is that smolagents is the only adapter here that survives three
consecutive 429s - and it does so by blocking for two to four minutes on a
single item. Measured five times at 139 s, 160 s, 213 s, 220 s and 225 s, all
inside the bracket the constants predict.
Setting Retry-After: 2 shortens the first two gaps to exactly 2.0 s and leaves
the third at 213 s, which is what pins the delay on the outer layer rather than
the client. Nothing the provider says can shorten it; only passing
retry=False, or a smaller RETRY_WAIT, would.
For a batch job that is a throughput collapse the scorecard cannot see, because the item passes. For an interactive script it is arguably the right default. See transport.md.
It applies to rate limits and nothing else. Retrying is constructed with
retry_predicate=is_rate_limit_error, which lowercases the exception text and
looks for 429, rate limit, too many requests or rate_limit. A hung
provider produces none of those, so the outer layer never engages: against a
20-second hang this adapter gives up in 8.5 s across 3 attempts, which is the
OpenAI client's own two retries and nothing more. "smolagents retries for
minutes" is true of exactly one status code.
One consequence worth noting: because the predicate matches on the string, a provider whose 500 body happens to mention rate limiting would trigger the minutes-long path too.
Timeouts reach it only through client_kwargs. OpenAIServerModel builds
its own client and takes no timeout of its own, so the adapter passes
client_kwargs={"timeout": config.request_timeout_s}. Without that line the
shared budget is silently ignored and the OpenAI client's ten-minute default
applies - which is what it was doing until
the parity fix.
Not yet done¶
- No pause support.
smolagentshas no interrupt/approval primitive equivalent to LangGraph'sinterrupt()or the OpenAI SDK'sneeds_approval, so the adapter implements noresumeand the two pause arenas report unsupported rather than failed. Emulating one by hand (asvanilladoes) would measure the adapter, not the framework. CodeAgentruns the three single-agent tool arenas. It ships as its own entry,frameworks/smolagents_code— a genuinely different execution model, so a separate adapter rather than a swap inside this one. It is scoped totool_use,ragandstructured_output, where theToolCallingAgentcontrast is meaningful;pause,durableandmulti_agentare out of scope for the same reasons they are here.