Choose a framework or SDK¶
A framework supplies building blocks inside an application you design. Use this route when you are choosing orchestration, tool, state, approval, or delegation primitives. If you want a complete environment that already owns prompts, execution, permissions, persistence, and interaction, use the coding-harness route.
Coverage counts below refer to the original seven arenas. The eighth, boundary response, currently covers only vanilla, LangGraph and OpenAI Agents SDK.
One page per adapter — wiring notes, the gotchas that cost real debugging time, and results. Written and maintained by whoever owns the adapter.
| Framework | Status | The thing worth knowing |
|---|---|---|
| LangGraph | runs all 7 arenas | Ties the hand-rolled baseline on the wire, byte for byte; native interrupt() + on-disk checkpointer; loses res-01 on malformed tool args |
| OpenAI Agents SDK | runs all 7 arenas | Heaviest of the in-band six on the wire (1.14×); needs_approval + a fully serialisable RunState; tracing uploads to OpenAI unless disabled |
| Pydantic AI | runs all 7, green on all 7 | Agent(retries=...) is not a loop cap — it ran 50 LLM calls on a budget of 6; deferred tools for the pause; its hand-built delegation chain costs 2N without the library having a delegation feature |
| Microsoft Agent Framework | runs 6 of 7 | Tool loop uncapped by default (41 calls on a budget of 6); pauses natively via approval_mode, but the pause dies with the process |
| Google ADK | runs all 7 | The only real loop cap out of the box (N means N); needs litellm to leave Google, the heaviest dep tree here; loses both res-01 and res-02 to uncaught exceptions |
| smolagents | runs 5 of 7 | 3.90× baseline on the wire — a 4.2 KB templated system prompt resent every request; recovers 8/8 scripted faults, with 3× cost on four validation failures |
| CrewAI | not in CI | Drives a text ReAct loop, not native tool calling — answers correctly, records no tool calls |
| Claude Agent SDK | stub, on purpose | Spawns the claude CLI over the Anthropic Messages API; cannot sit behind the shared OpenAI-compatible gateway |
The dependency-free vanilla baseline is
documented next to its code. It is the control in the experiment, and it is
not the cheapest on the wire — see overhead.md.
vanilla, pydantic_ai, and smolagents are green on every arena they run.
langgraph and openai_agents each lose one resilience item, while
google_adk loses two. Those are measured findings, not broken adapters.
microsoft_af pauses 12/12 but is unsupported on durable_state; smolagents
is unsupported on both. Reported as unsupported rather than failed.
Reading these pages¶
Every profile starts with the same decision and evidence summary: best fit, owned responsibilities, important limitation, evidence status, pinned version, and review deadline. Read the profile evidence guide before treating a source review or mock result as a product ranking.
Pass rates in mock mode are ~100% by construction and prove only that an adapter
is wired correctly. The columns that compare frameworks honestly are marked
(comparable) on each page, and collected by python -m arena summary.