Skip to content
Hammad Yousuf.
LLM streaming latency · NOTE 17

LLM streaming latency: measure first token and finish

The first word arrived quickly. The useful answer still took too long.

Starter template · 4 October 2026
Hammad Yousuf
By Hammad YousufSoftware, AI & automation
FIRST TOKEN → TOOL WORK → ACCEPTED FINISH
Streaming measurement sheet · Open access

What to inspect

Measure time to first visible output separately from time to a complete accepted result. Streaming can improve perceived responsiveness while the tool work or validation still takes time. Observe cancellation and interrupted responses as well as normal completion.

Worked example

Illustrative timeline: the first token appears at 0.5 seconds, a tool starts at 2 seconds and the accepted answer arrives at 9 seconds. The user experienced both a fast start and a longer wait. These timings are fictional, used to show what to measure.

Copy the full prompt

Adapt the inputs to your task. Remove private data before sending anything to a model. This page copies text locally; it does not run the prompt.

Full prompt · adapt the inputs
Create an instrumentation plan for [STREAMING WORKFLOW]. Record request start, first visible token, tool start/end, response completion and acceptance completion. Include a cancellation and an interrupted stream. Distinguish observed latency from placeholder values. Specify how partial output is presented so it cannot be mistaken for a verified final result.

Review checklist

  1. Record first visible token and accepted completion separately.
  2. Include tool and validation time.
  3. Test cancellation and interrupted streams.
  4. Label partial results visibly.

Keep the resource

Download the complete note ↓Markdown · explanation, example, prompt and checklistDownload the prompt ↓Plain text · ready to adaptOpen the full-size visual ↓SVG · scalable reference diagram

Documentation & scope

Anthropic: streaming ↗

This is an original workflow template with an illustrative example. It does not report completed model tests or measured performance. Provider APIs can change; check the linked documentation for your exact integration.

Have a workflow to review?

Bring the task, inputs and the result you need. We can discuss the integration and the checks around it.

Discuss your workflow ↗