Insights

Field notes from building and operating AI systems in production: what breaks, why it happens, and how to test for it.

Evaluating models 10

Qualifying models and building evals and tests you can trust.

  1. We made five model picks in one day Four ad hoc tests gave four model picks, and a fifth fell to our own scoring. The single replay protocol we use now, and the order we read its results in.
  2. Half our replay sample tested the wrong prompt We rebuilt production requests for a model eval and paired nearly half with answers to a different prompt. How it showed, what we withdrew, and the guards.
  3. What broke when we tested 20+ models on one tool call We ran more than twenty models through one structured-decision task with a forced tool call. About half failed on liveness. The failure types, and what passed.
  4. Recording the bad day: replay tapes for rate limits and retries Our record and replay setup could not hold a 429 followed by a 200. What a sequence recorder needs, and how to capture failures no host gives on demand.
  5. Testing tool calling on Hugging Face Inference Providers A method for qualifying tool calling across Hugging Face Inference Providers: the endpoint, provider pinning, fixture design and the traps. No measured results.
  6. A stub LLM that always returns clean JSON hides your parsing bug Our tests passed because the stub returned tidy JSON and the check only looked at the status code. Real models add fences and prose. Test the dirty case.
  7. Record and replay for LLM tests Record against live APIs once, replay in CI for fast, cheap, deterministic tests. Plus a static test that stops placeholder values reaching the model as facts.
  8. Using an independent LLM judge, and its limits A method note on LLM-as-judge: keep the judge independent, keep samples honest, treat a past production answer as a baseline and not as truth.
  9. Grade the tool call, not the text Our first benchmark run scored zero for every model because the grader read message text while the answer was in tool-call arguments. How to grade it right.
  10. How to qualify an LLM router model: hard gates before cost Sorting candidate models by price first surfaces broken ones. Disqualify on failures, then rank on cost and latency. The gates, the process and the traps.