Choose what you are comparing¶
Agentic Arena compares two different kinds of developer choice. Pick the system boundary first; the relevant evidence and methodology follow from it.
Choose a framework or SDK
You are building an agent application and need orchestration, tools, state, approvals, or delegation primitives.
Compare frameworks → Complete developer systemChoose a coding agent or harness
You need an environment that already owns prompts, tools, execution, permissions, persistence, and user interaction.
Compare coding harnesses →The distinction in sixty seconds¶
| Framework or SDK | Coding agent or harness | |
|---|---|---|
| What you receive | Primitives used inside an application you design | A working developer experience around a model |
| Usually owns | Orchestration, tool registration, state APIs, handoffs, tracing | Prompts, context selection, tools, execution, permissions, sessions, UI, recovery |
| You still choose | Product UX, execution environment, policy, storage, most operational behavior | Model, workspace, authority, configuration, and operating environment |
| Examples here | LangGraph, Pydantic AI, OpenAI Agents SDK, Google ADK | OpenHands, OpenCode, Cline, goose, SWE-agent, Aider |
| Fair comparison | Hold model, tools, tasks, datasets, and scorer fixed | Hold task repository, model, runtime, authority, budgets, and final-state checks fixed |
Do not put both kinds of system in one score table. A complete harness owns many variables that the framework experiment deliberately holds fixed.
Start with the decision, then inspect the evidence¶
Apply your constraints
Turn deployment, control, recovery, and dependency requirements into a shortlist.
Read measured findings
See what Agentic Arena has actually observed and the command that regenerates it.
Check whether claims compare
Build two evidence records and expose mismatched modes, models, datasets, or scorers.
Audit the method
Inspect controls, unsupported states, scoring, and the boundary of every claim.
Framework questions¶
| Your question | Open |
|---|---|
| Does a framework support a capability? | Feature matrix |
| How much machinery does it add to the same task? | Framework overhead |
| How does it behave when a provider fails? | Transport failures |
| What does delegation cost? | Multi-agent comparison |
| How are tool contracts represented? | Tool schemas |
Coding-harness questions¶
| Your question | Open |
|---|---|
| Which complete systems belong in the comparison? | Coding-harness overview |
| How do their system boundaries differ? | Coding-harness matrix |
| What evidence exists for a specific harness? | Open its profile and read the evidence banner first |
| How will Agentic Arena run them fairly? | Harness admission and experiment contract |
Evidence behind both routes¶
Evaluation arenas define the workload and scorer. The fairness controls state what must remain fixed. The profile evidence guide explains source reviews, diagnostics, mock tests, native-live runs, freshness, and independent reproduction.