Skip to content
Hammad Yousuf.
LLM evaluation · NOTE 01

LLM evaluation: measure the usable result

The fastest LLM can still finish the job last.

Starter template · 4 October 2026
Hammad Yousuf
By Hammad YousufSoftware, AI & automation
SAME INPUT → ALL ATTEMPTS → ACCEPTED RESULT
Model comparison scorecard · Open access

What to inspect

Choose one task and define acceptance before comparing models. Measure the complete workflow, including tool calls, rejected outputs and retries. Keep the input cases and tool access the same, and repeat trials to expose variation.

Worked example

Illustrative arithmetic: workflow A costs $0.02 per attempt and needs three attempts for one accepted result. Workflow B costs $0.04 and succeeds once. The accepted results cost $0.06 and $0.04 respectively. These are invented inputs to explain the calculation, not measured model prices.

Copy the full prompt

Adapt the inputs to your task. Remove private data before sending anything to a model. This page copies text locally; it does not run the prompt.

Full prompt · adapt the inputs
Design an evaluation for [TASK] using [INPUT CASES]. Define an explicit pass/fail rule for each case before running anything. Record model version, prompt version, tool access, attempt count, total cost and time to an accepted result. Include all failed attempts. Separate measured values from unknown values. Return a blank scorecard and a repeatable run protocol. Do not choose a winner without completed trials.

Review checklist

  1. Freeze the input cases and acceptance rules.
  2. Record model, prompt and tool versions.
  3. Count retries and failed attempts in cost and elapsed time.
  4. Repeat trials and report spread as well as the average.

Keep the resource

Download the complete note ↓Markdown · explanation, example, prompt and checklistDownload the prompt ↓Plain text · ready to adaptOpen the full-size visual ↓SVG · scalable reference diagram

Documentation & scope

OpenAI: evaluation best practices ↗

This is an original workflow template with an illustrative example. It does not report completed model tests or measured performance. Provider APIs can change; check the linked documentation for your exact integration.

Have a workflow to review?

Bring the task, inputs and the result you need. We can discuss the integration and the checks around it.

Discuss your workflow ↗