Skip to content

Problem register

Start with the symptom, reproduce it, then select a remedy. These entries connect existing framework evidence to portable harness requirements. Results apply to the published configurations, not every version or use of a library.

Worked investigation: a costly repeating loop

  1. Capture request count and tool-call IDs at the model boundary. Do not begin with prompt tuning.
  2. If requests repeat with identical context, check error delivery. A model cannot correct an error it never sees.
  3. If requests multiply during gateway failures, inspect retry ownership and distinguish attempt timeout from task deadline.
  4. If each request grows, measure context accumulation; separate its slope from constant schema/prompt overhead.
  5. If roles multiply the bill, account for delegation, including copied context and advertised schemas.
  6. Change one design variable and rerun the same fixture. Check useful task completion as well as lower cost.

A correction is successful only when the original failure is absent and allowed work still completes. Report configuration, model mode, versions, missing adapter coverage, and side effects. Never translate a synthetic token estimate into an unqualified production bill.

New entries should contain symptom, minimal reproducer, affected configuration, evidence, alternatives, uncertainty, and an acceptance experiment. See the experiment template.