Next phases — research notes and plan¶
Historical plan. The current concept-first ecosystem direction is described in the charter and ordered backlog. The notes below retain the rationale for the earlier framework-comparison program. Statements about missing capabilities describe that earlier snapshot; consult current source, findings, and the backlog before acting on them.
Written 2026-09-01. Companion to ROADMAP.md; ROADMAP stays the terse index, this file records the reasoning and the current batch of work.
Status update. Tracks A, B and C below are done, and so is everything the 2026-09-01 note put "out of scope for this batch" bar the three standing blockers. All seven arenas run, seven adapters plus five
*_multipipeline entries are mock-green, delegation cost is measured to an exact law, the durable pause runs across a real process restart, and every arena-owned control is enumerated with the test that holds each adapter to it (fairness-controls.md). Since then an eighth entry —smolagents_code, aCodeAgentcontrast to theToolCallingAgentsmolagentsadapter — runstool_use/rag/structured_outputat 6.95× baseline wire cost, andarena validategrew a reachability check for a tool/pause check its mock scenario can never satisfy.python -m arena chartnow renders the scorecard as SVG bar charts (Phase 3's "latency / token / cost charts" item). The reasoning below is kept as the record of how that batch was planned.
Where things stand¶
Phases 0–3 are done: seven arenas, seven adapters plus five pipeline entries all run against the mock, CI is green, and the harness core is still stdlib-only with zero runtime dependencies. What the repo has actually measured is collected in findings.md.
The remaining gaps are no longer breadth. They are: no live scorecard (needs an
API key — every number published so far is mock mode, so nothing says whether any
framework gives better answers), crewai installs and answers on Python 3.12
but its tool_used evidence is not yet confirmed (a step_callback capture path
was added; needs a debug-workflow run), and claude_agent_sdk is still a stub.
What the landscape says (Sept 2026)¶
The Python agent-framework field has settled into two groups:
- Provider-native SDKs — OpenAI Agents SDK, Claude Agent SDK, Google ADK. Optimised for one model family.
- Provider-neutral frameworks — LangGraph, CrewAI, Pydantic AI, Microsoft Agent Framework (the merged AutoGen + Semantic Kernel line), Smolagents.
Public comparisons still lean on prose and one-off token counts (e.g. "~18% token overhead vs LangGraph", "$390 vs $1088 over 90 days"). Nobody publishes a regenerable cross-framework scorecard on identical tasks — which is the gap this project fills. The takeaway for us: keep the harness honest and reproducible, and widen coverage.
Installability check (this machine: Windows, CPython 3.14.7, no C compiler)¶
Ran pip install resolution for each stub package into a clean 3.14 venv:
| Package | 3.14 result | Notes |
|---|---|---|
pydantic-ai-slim[openai] |
✅ installs + imports | 2.37.0; all cp314 wheels; native typed output |
openai-agents |
✅ installs + imports | 0.22.0; needs tracing disabled offline |
agent-framework (meta) |
⚠️ huge tree | pulls agent-framework-core[all] (azure, boto3, redis, qdrant, numpy…). Use the narrow agent-framework-core + agent-framework-openai instead |
claude-agent-sdk |
✅ installs | but see below — wrong protocol shape |
crewai |
❌ (unchanged) | chromadb/onnxruntime have no 3.14 wheels; 3.12-only, still unverified |
claude-agent-sdk installs but does not fit the harness as-is: it drives the
claude CLI (Node) as a subprocess and speaks the Anthropic Messages API, not an
OpenAI-compatible /chat/completions. It cannot point at the OpenAI-shaped mock
server without either an Anthropic-shaped mock or a translating proxy (LiteLLM) and
Node in CI. Left as a stub with that blocker written down.
This batch of work¶
Branch: next-phases. One reviewable PR, small separate commits, no merge to
main. Every commit keeps ruff + pytest green; every new adapter is proven
15/15 against the mock before its commit lands.
Track A — framework adapters (Phase 2)¶
pydantic_ai—OpenAIChatModel+OpenAIProvider(base_url, api_key), shared tools via@agent.tool_plain, tool calls read fromresult.all_messages(), tokens from the run usage. Pinpydantic-ai-slim.openai_agents—AsyncOpenAI(base_url, api_key)+OpenAIChatCompletionsModel+Agent/Runner,set_tracing_disabled(True),function_toolwrappers, usage + tool calls from theRunResult. Pinopenai-agents.microsoft_af— attempt with the narrowagent-framework-openaiinstall (OpenAIChatClient+ChatAgent). Ship only if it installs lean and runs 15/15; otherwise leave an improved stub noting the dependency weight.claude_agent_sdk— stays a stub; add a README with the protocol-mismatch blocker and the two ways a contributor could close it.
Track B — second arena (Phase 3): structured_output¶
- New mechanical scorer checks, stdlib-only, unit-tested:
not_contains,json_valid,json_path_equals(tolerance-aware), and a tiny inlinejson_schemavalidator (object / required / properties / type / minItems / additionalProperties — enough for the design schema, no new dependency). arenas/structured_output/—arena.toml, ~15-itemdataset.jsonlover an expanded corpus,mock_script.jsonwith valid JSON per item plus one bad-JSON-then-retry scenario.- Grow
arena/tools/corpus.jsonfrom 10 to ~28 passages (also unblocks the futureragarena).tool_usedataset/mock untouched.
Track C — wiring¶
ci.yml: addmock-smokematrix rows for each newly verified adapter.- Exact version pins in
frameworks/<name>/requirements.txt. - Fill the
❓cells indocs/feature-matrix.mdfrom first-hand adapter work; turn thedocs/frameworks/*.mdstubs into real notes for the ones now built. - Update
README.mdtables,ROADMAP.md,STATUS.md,docs/arenas/README.md. results/is untouched — it only ever holds live numbers, and there is no key wired in. Producing the first live scorecard stays a separate, key-gated task.
Out of scope for this batch¶
- ~~
rag,human_in_the_loop,durable_statearenas~~ — all shipped since, with the harness resume/checkpoint API (ResumableRunner, leg merging) they needed. crewaiverification — still needs a 3.12 environment; the adapter now installs and answers there, tool-call capture pending a debug run.- Live scorecards (need
OPENAI_API_KEYin thefull-runworkflow). - MkDocs site (Phase 4) —
mkdocs-materialbuilds but not under--strict: the docs deep-link to source files outsidedocs/(legitimate for GitHub browsing), so a Pages build would need those rewritten or a relaxed config.