Scorecard format example¶
Not real numbers. This is a
--mode mockrun of thetool_usearena, committed only to show the shape of a generated scorecard. Mock pass rates are ~100% by construction and the token/latency figures are client-serialisation artifacts (see methodology §5). Never cite these.
Real scorecards are generated into results/<arena>/ by a --mode live run.
A single agent must answer factual and arithmetic questions using a search tool and a calculator tool. Tests the core tool-calling loop and whether the framework feeds tool results back to the model correctly.
- Mode: mock ⚠️ plumbing only — not a quality signal
- Model: mock-model
- Dataset: 15 items × 1 repeat(s)
- Run at: 2026-09-01T03:10:31Z (0.76s)
- Harness: v0.1.0 · Python 3.14.7
| Framework | Ver | Pass rate | Errors | Mean latency | Mean tokens | Mean LLM calls | Est. cost |
|---|---|---|---|---|---|---|---|
vanilla |
stdlib (arena 0.1.0) | 100% (15/15) | 0 | 0.018s | 448 | 2.07 | $0.0035 |
Generated by python -m arena scorecard. Do not edit by hand.