Skip to content

Scorecard format example

Not real numbers. This is a --mode mock run of the tool_use arena, committed only to show the shape of a generated scorecard. Mock pass rates are ~100% by construction and the token/latency figures are client-serialisation artifacts (see methodology §5). Never cite these.

Real scorecards are generated into results/<arena>/ by a --mode live run.


A single agent must answer factual and arithmetic questions using a search tool and a calculator tool. Tests the core tool-calling loop and whether the framework feeds tool results back to the model correctly.

  • Mode: mock ⚠️ plumbing only — not a quality signal
  • Model: mock-model
  • Dataset: 15 items × 1 repeat(s)
  • Run at: 2026-09-01T03:10:31Z (0.76s)
  • Harness: v0.1.0 · Python 3.14.7
Framework Ver Pass rate Errors Mean latency Mean tokens Mean LLM calls Est. cost
vanilla stdlib (arena 0.1.0) 100% (15/15) 0 0.018s 448 2.07 $0.0035

Generated by python -m arena scorecard. Do not edit by hand.