Dependencies, pins, and deprecations¶
The policy¶
- The harness (
arena/) has zero runtime dependencies. Standard library only, including the mock LLM server, the JSON-schema checker, and the HTTP client. A PR that adds a runtime dependency toarena/needs a very good reason, because the harness is what every framework is measured through — its own footprint should not be part of the comparison. devisruff+pytest. Deliberately no frameworks: a plain test job must stay fast and installable everywhere. The consequence is that the wire-level contract tests can only seevanillathere, which is why they also run in thecomparisonCI job — see methodology.md §4.ruffis pinned exactly;pytestis floored.ruff format --check .is a CI gate, and both the formatter's output and the rules behindselect = [E, F, I, UP, B, SIM]change across releases — a range lets a contributor's localruff formatdisagree with CI and produce a diff nobody asked for.pytesthas no such gate: the tests pass or they don't.- Every adapter pins exactly. A scorecard records the library version it was
produced with, so a bump invalidates that scorecard until it is re-run. Ranges
are used only for adapters that are not yet verified (
crewai).
Current pins¶
| adapter | pins |
|---|---|
vanilla |
none — stdlib |
langgraph |
langgraph==1.2.11, langchain-core==1.6.1, langchain-openai==1.6.0 |
pydantic_ai |
pydantic-ai-slim[openai]==2.37.0 |
openai_agents |
openai-agents==0.22.0 |
microsoft_af |
agent-framework-core==1.16.0, agent-framework-openai==1.14.1 |
smolagents |
smolagents[openai]==1.26.0 |
google_adk |
google-adk==2.8.0, litellm==1.99.0 — LiteLLM is required, not optional: it is the only way ADK reaches a non-Google endpoint, and it is the heaviest dependency tree in the repo |
openai_agents_multi |
-r ../openai_agents/requirements.txt — same reason as langgraph_multi |
langgraph_multi |
-r ../langgraph/requirements.txt — the pipeline contrast entry shares the single-agent adapter's pins, so the two cannot drift apart |
vanilla_multi |
none — stdlib |
crewai |
crewai>=0.130,<1.0 — a range, because the adapter is not yet verified |
claude_agent_sdk |
unpinned — deliberate stub |
Two adapters deliberately install the narrow package rather than the
meta-package: pydantic-ai-slim[openai] instead of pydantic-ai, and
agent-framework-core + -openai instead of agent-framework. The
meta-packages pull provider SDKs (azure, boto3, redis, qdrant, ollama, ...) that
no arena uses.
smolagents runs the opposite way: the bare package is too narrow. Without the
[openai] extra, OpenAIServerModel raises at construction, and the failure
reads like a bad import rather than a missing extra. This matters for install time and for honestly describing what a
framework costs to adopt for this task.
Automation¶
.github/dependabot.yml opens one grouped PR per adapter, monthly. The cadence is
deliberately slow: the point is to make drift visible on a predictable schedule,
not to keep main on latest. A bump PR is reviewed by running the contract tests
and the mock sweep, and by checking the comparison job's overhead table — a jump
there means the library changed how it serialises tool schemas, which is itself a
finding worth recording in overhead.md.
What CI does and does not exercise¶
A green CI run on a bump PR is not the same as "this bump is verified". The gap worth knowing about:
| action | used in | exercised by CI? |
|---|---|---|
actions/checkout |
every workflow | yes |
actions/setup-python |
every workflow | yes |
actions/upload-artifact |
full-run.yml, and the comparison job |
yes |
upload-artifact used to appear only in full-run.yml, which is
workflow_dispatch and needs OPENAI_API_KEY — so a bump to it landed untested
and would first have run whenever someone triggered a live run. The comparison
job now uploads the cross-arena summary with the same action, which closes that
gap incidentally. Worth re-checking this table whenever a workflow is added.
actions/checkout@v7 carries a breaking change — it blocks checking out fork PRs
under pull_request_target and workflow_run. This repo's workflows trigger on
push, pull_request and workflow_dispatch only, so it does not apply. Worth
re-checking if a workflow ever adopts one of those triggers.
Deprecation register¶
Known upstream deprecations that affect an adapter, with the decision made about each. An entry stays here until the adapter no longer triggers it.
langgraph.prebuilt.create_react_agent¶
LangGraphDeprecatedSinceV10: create_react_agent has been moved to
`langchain.agents`. Please update your import to
`from langchain.agents import create_agent`.
Deprecated in LangGraph V1.0 to be removed in V2.0.
- Where:
frameworks/langgraph/adapter.py - Status: not migrated, deliberately.
- Why not: the replacement lives in the
langchainpackage, which the adapter does not currently install — it depends onlangchain-coreonly.langchainis also on a separate version track (1.3.x stable) from the pinnedlangchain-core(1.6.1), so adopting it means resolving a new top-level dependency against existing exact pins. LangGraph 2.0 is not released, andcreate_react_agentworks correctly on the pinned 1.2.11 — all five arenas are green and the adapter passes every wire-level contract test. - Trigger to revisit: LangGraph 2.0 reaching a release candidate, or a
Dependabot bump that brings
langchainin as a transitive dependency anyway. - Migration sketch: add
langchain==<compatible>toframeworks/langgraph/requirements.txt, swap the import tofrom langchain.agents import create_agent, and re-check the overhead table — a different agent constructor may serialise tool schemas differently, which would move LangGraph's position in overhead.md.
Recording the decision is the point. An unexplained deprecation warning in CI output is noise that everyone learns to scroll past; a dated entry with a trigger is something the next person can act on.