Skip to content

Feature matrix

Things that don't reduce to a scorecard number. Fill a cell only from first-hand experience writing the adapter โ€” link to the adapter code or an upstream doc.

Legend: โœ… built-in ยท ๐ŸŸก possible with work ยท โŒ not really ยท โ“ not yet assessed

Capability vanilla LangGraph CrewAI OpenAI Agents SDK Claude Agent SDK Pydantic AI MS Agent Framework smolagents Google ADK
Language Python Python Python Python Python Python Python / .NET Python Python
OpenAI-compatible base_url โœ… โœ… โœ… (LiteLLM) โœ… โ“ โœ… โœ… (Chat Completions client) โœ… (OpenAIServerModel, needs the [openai] extra) โœ… (via LiteLLM only)
Streaming tokens โŒ โœ… ๐ŸŸก โ“ โ“ โ“ โ“ โ“ โ“
Recovers from malformed tool args โœ… โŒ โœ… โœ… โ“ โœ… โœ… โœ… โŒ (raises)
Runs every tool call in a batched turn โœ… ๐ŸŸก (drops a malformed sibling) โŒ (one Action: per turn โ€” text ReAct) โœ… โ“ โœ… โœ… ๐ŸŸก (drops the whole batch) โœ…
Recovers from an unknown tool name โœ… โœ… โœ… โŒ (raises) โ“ โœ… โœ… โŒ (not written back) โŒ (raises)
Native OpenAI tool calling โœ… โœ… โŒ (text ReAct loop) โœ… โ“ โœ… โœ… โœ… (plus a final_answer control tool) โœ… (through LiteLLM)
Tool-call history exposed โœ… โœ… ๐ŸŸก (via wrapper) โœ… (new_items) โ“ โœ… (all_messages()) โœ… (messages contents) โœ… (memory.steps) โœ… (event stream)
Token usage exposed โœ… โœ… (usage_metadata) โœ… (usage_metrics) โœ… (context_wrapper.usage) โ“ โœ… (result.usage) โœ… (usage_details) โœ… (per-step token_usage) โœ… (per-event usage_metadata)
Built-in multi-agent โŒ โœ… (graph, measured) โœ… (crew) โœ… (handoffs, measured) ๐ŸŸก (subagents) ๐ŸŸก โœ… โœ… (managed_agents, measured) โœ… (sub_agents and AgentTool, both measured)
Human-in-the-loop / interrupts ๐ŸŸก (emulated, measured) โœ… (interrupt, measured) โ“ โœ… (needs_approval, measured) โ“ โœ… (deferred tools, measured) โœ… (approval_mode, measured) โŒ (no interrupt primitive) ๐ŸŸก (reported, not enforced, measured)
Durable state / checkpointing ๐ŸŸก (stateless resume, measured) โœ… (SqliteSaver, measured) โ“ โœ… (RunState.to_json, measured) โ“ ๐ŸŸก (stateless resume, measured) โŒ (session store does not round-trip, measured) โŒ โœ… (DatabaseSessionService, measured)
Typed / schema-validated output ๐ŸŸก ๐ŸŸก ๐ŸŸก ๐ŸŸก (output_type) โ“ โœ… ๐ŸŸก ๐ŸŸก ๐ŸŸก (output_schema)
Async API โŒ โœ… ๐ŸŸก โœ… (run_sync wraps it) โ“ โœ… โœ… (async-only) ๐ŸŸก (arun) โœ… (async-first)
Observability hooks / tracing โŒ โœ… (LangSmith) โœ… (events) โœ… (built-in; disabled for the arena) โ“ ๐ŸŸก (Logfire) โœ… (OpenTelemetry) โœ… (OpenTelemetry) โœ… (OpenTelemetry)
Retries a transient 429 โŒ (measured) โœ… 1 retry (measured) โ“ โœ… 2 retries (measured) โ“ โœ… 2 retries (measured) โœ… 2 retries (measured) โœ… retries until it wins, +2โ€“4 min (measured) โœ… 2 retries (measured)
Licence โ€” MIT MIT MIT โ“ MIT MIT Apache-2.0 Apache-2.0

The two Recovers from ... rows are measured, not judged โ€” see the resilience arena and the comparison CI prints on every run.

The Typed / schema-validated output row is the opposite: a library-capability claim, not a measurement. On the structured_output arena every adapter โ€” pydantic_ai included, with its output_type=str โ€” drives the model with a prompt and sends no response_format, so nothing here has exercised a native mechanism. See structured-output.md; wiring one per framework is a planned batch of contrast entries, not a filled cell yet.

Retrying the gateway

The Retries a transient 429 row is measured by scripting the provider failing rather than the model โ€” see transport.md. It is the one dimension where the hand-rolled baseline loses outright: vanilla has no retry at all, so a single 429 loses the item, while every framework survives one.

Two differences worth knowing. langgraph retries once where the others retry twice, giving up an attempt earlier against a provider that rate-limits in bursts. And smolagents is the only one that survives three consecutive 429s โ€” by sleeping for two to four minutes on a single item, which the scorecard cannot see because the item passes. That sleep is rate-limits only: against a provider that simply hangs, smolagents fails as fast as everyone else.

A hung provider is a separate row from a failing one, and it is the arena's own request_timeout_s that bounds it โ€” a knob five of seven adapters were ignoring until it was wired through and gated. The budget bounds an attempt, not an item, so a framework that retries twice can spend three times it.

Batched tool calls, and a quieter failure mode

A model may return several tool calls in one turn. With two valid calls all seven adapters do the right thing โ€” both run, both results reach the model. (crewai is absent: it drives a text ReAct loop and emits one Action: per turn, so a batched turn never arises.) The interesting case is when one call in the batch is broken. For each fault, batched with one good call, two questions of the next request: did the successful call's result reach the model, and was the broken call reported at all?

unknown tool malformed args missing required arg
vanilla both both both
pydantic_ai both both both
microsoft_af both both both
langgraph both good only both
smolagents error only good only error only
openai_agents raises both both
google_adk raises raises both

Three distinct ways to mishandle a batch, and they are not equally bad:

  • Silent partial (langgraph, and smolagents on malformed args). The broken call vanishes with no message of any kind. The run continues, answers from partial evidence, and the model is never told a call went missing. This is the quietest failure here and the one worth watching for.
  • The successful sibling is discarded (smolagents, two faults). The error is surfaced โ€” but as a rewritten task (New task: โ€ฆ Error: โ€ฆ Now let's retry) rather than as a turn in the transcript, and the good call's observation is dropped along with the history. The model then re-emits the identical batch, because from its point of view it never ran anything. This is the same mechanism as its resilience losses, seen from a different angle.
  • Raises (openai_agents, google_adk). Loud, and the same root cause as those frameworks' resilience losses. Nothing is silently wrong, which makes it the best of the three.

That is a different question from the one resilience asks: not "does it recover?" but "does it tell the truth about what happened on the way?" A framework that drops a lookup without saying so produces a confident answer built on half the evidence.

It also refines the res-01 row. langgraph losing malformed tool arguments is conditional: alone, the malformed call produces no tool message and the graph halts (so the item fails, visibly); batched with a call that succeeds, the graph carries on and the drop becomes invisible. The visible failure is the better of the two outcomes.

Measured in tests/test_parallel_tool_calls.py, which gates the invariants (valid batches work, the baseline surfaces every result, batching never makes a framework worse than it is serially) and leaves the per-framework differences as findings, the same way resilience does.

An earlier version of this table read "missing required arg โ€” smolagents 0 of 2, the whole batch is dropped". That was measured with a helper that counted the word "observation" in the returned messages โ€” and a corpus entry describes Tokyo Tower as a "communications and observation tower", so it read 3 for a batch of 2. The direction was right and the zero was real, but "the model is never told" was wrong for smolagents: it is told, in a rewritten task, and what it loses is the successful sibling. The probe now matches each outcome's own text, and a test pins that none of those markers appear in the corpus.

Those two rows do not fully capture smolagents, whose real boundary on resilience is its tool-validation layer: any failure raised before the tool body runs (unknown name, missing argument, unexpected argument, null arguments) is never written back into the conversation, so the model cannot see it and repeats the identical call until the step budget is gone. It still answers each one from its last memory step โ€” 8/8 โ€” but at six LLM calls against the two the other four faults take. See smolagents.md.

The Built-in multi-agent row is now partly measured. multi_agent carries two real three-role pipelines alongside the single-agent entries:

comparison prompt LLM calls
vanilla -> vanilla_multi (hand-rolled pipeline) 2.50x 2.00x
langgraph -> langgraph_multi (StateGraph) 2.50x 2.00x
vanilla_multi -> langgraph_multi (graph machinery alone) 1.00x 1.00x
openai_agents -> openai_agents_multi (native handoffs) 2.64x 2.00x
smolagents -> smolagents_multi (managed_agents) 3.93x 3.00x

The cost of multi-agent is the structure, not the framework: three roles double the LLM calls and 2.5x the prompt tokens whether you build them with a graph library or a for loop, and LangGraph's orchestration adds nothing on top โ€” vanilla_multi and langgraph_multi send the byte-identical request. (It read 2.62x / 0.97x until a LangChain release equalised the tool-schema difference from overhead.md; report_delegation.py gates the ratios now.)

This measures cost with benefit held at zero โ€” the mock scripts identical turns, so all four entries return the same brief. Whether delegation improves the answer needs a live run. See multi-agent.md.

Model-decided delegation (a native handoffs chain) costs ~10% more prompt than the structural pipelines at the same call count, and 94% of that difference is the transfer_to_* schemas riding on every request rather than the transfers themselves โ€” you pay for a handoff by advertising it, not by taking it. A supervisor offering N handoffs pays for N schemas on every request in the run.

A sub-agent invoked as a tool (managed_agents, AgentTool) is a third shape, and it is the only one that costs 3x the model calls rather than 2x: its reply is a tool result, not the end of the run, so every delegator has to wake up and produce its own final answer afterwards.

Measured from one role to five across five implementations, each follows an exact law - handoffs N+1, ADK sub_agents N+2, sub-agent-as-tool 2N - and the 2N law replicates in three libraries that share no code, so it is a property of the mechanism. The third replication is the decisive one: Pydantic AI has no delegation feature, so its chain is an ordinary async tool awaiting a nested run, and it costs 2N anyway. But the call count is only half the bill: a transfer keeps one growing conversation (prompt compounds 15.4x over four roles) while a sub-agent starts fresh (3.59x ADK, 3.68x pydantic_ai). Inside ADK, AgentTool costs a third more calls and a quarter of the prompt than sub_agents. See multi-agent.md.

Still judged: CrewAI crews, the same sub-agent-as-tool shape, blocked on an adapter that has never been mock-verified.

The Human-in-the-loop / interrupts row is now measured for six adapters. human_in_the_loop observes the pause in the harness rather than trusting the agent's prose:

adapter result mechanism
langgraph 12/12 native โ€” interrupt() + MemorySaver, resumed with Command(resume=...)
openai_agents 12/12 native โ€” the tool is needs_approval=True, the run stops with a ToolApprovalItem, resume calls approve/reject
pydantic_ai 12/12 native โ€” the tool raises CallDeferred, the run returns DeferredToolRequests, resume passes deferred_tool_results
microsoft_af 12/12 native โ€” @tool(approval_mode="always_require") + ToolApprovalMiddleware; resume answers with to_function_approval_response
google_adk 12/12 LongRunningFunctionTool โ€” the call is reported in long_running_tool_ids, but the run does not stop on its own (hence ๐ŸŸก); resumed with a FunctionResponse against the same call id
vanilla 12/12 emulated โ€” transcript carried back in, no checkpoint (hence ๐ŸŸก, not โœ…)

Both produce an identical trace to the scorer: one pause per item, book_room never called before it, and called on exactly the six approved items.

Durable state / checkpointing is measured by durable_state, which discards the runner at the pause and rebuilds it, so only a real checkpoint or a serialised transcript can get across:

adapter result mechanism
langgraph 8/8 SqliteSaver writing to the harness-owned checkpoint dir; a fresh runner reopens the same thread
openai_agents 8/8 RunState.to_json() / from_json() โ€” the SDK serialises the whole run, not just the messages
pydantic_ai 8/8 stateless resume โ€” the conversation is serialised with ModelMessagesTypeAdapter and replayed as message_history
google_adk 8/8 DatabaseSessionService on sqlite+aiosqlite, writing to the harness-owned checkpoint dir; needs sqlalchemy on top of an already heavy tree
vanilla 8/8 stateless resume โ€” the whole transcript is serialised into resume_state. Durable, but it is not a checkpointer, hence ๐ŸŸก

An adapter patched to restart from scratch instead of resuming drops to 0/8: it reaches the right answer by redoing both lookups, and the call_counts check catches the duplicated work.

microsoft_af pauses (12/12) but is unsupported on durable_state, and that was measured rather than assumed: its approval state serialises cleanly, but the conversation lives in the AgentSession's in-memory store, which comes back from a JSON round trip as raw strings โ€” and restoring the approval state into a rebuilt agent re-queues the request instead of consuming the answer. The adapter therefore does not expose resume on a durable arena at all, because keeping one it cannot honour would post 0/8 and read as a broken framework.

crewai and smolagents have no resume method and report unsupported on both arenas. smolagents is marked โŒ rather than โ“ because it ships no interrupt or approval primitive to adapt: emulating one by hand, as vanilla does, would measure the adapter instead of the framework.

The โ“ cells are the backlog. Each one gets resolved when its adapter is written or when an arena exercises that capability directly.