# LLM streaming latency: measure first token and finish

Measure time to first visible output separately from time to a complete accepted result. Streaming can improve perceived responsiveness while the tool work or validation still takes time. Observe cancellation and interrupted responses as well as normal completion.

## Worked example

Illustrative timeline: the first token appears at 0.5 seconds, a tool starts at 2 seconds and the accepted answer arrives at 9 seconds. The user experienced both a fast start and a longer wait. These timings are fictional, used to show what to measure.

## Copyable prompt

```text
Create an instrumentation plan for [STREAMING WORKFLOW]. Record request start, first visible token, tool start/end, response completion and acceptance completion. Include a cancellation and an interrupted stream. Distinguish observed latency from placeholder values. Specify how partial output is presented so it cannot be mistaken for a verified final result.
```

## Checklist

- Record first visible token and accepted completion separately.
- Include tool and validation time.
- Test cancellation and interrupted streams.
- Label partial results visibly.

## Further reading

[Anthropic: streaming](https://platform.claude.com/docs/en/build-with-claude/streaming)
