What to inspect
Measure time to first visible output separately from time to a complete accepted result. Streaming can improve perceived responsiveness while the tool work or validation still takes time. Observe cancellation and interrupted responses as well as normal completion.
Worked example
Illustrative timeline: the first token appears at 0.5 seconds, a tool starts at 2 seconds and the accepted answer arrives at 9 seconds. The user experienced both a fast start and a longer wait. These timings are fictional, used to show what to measure.
Copy the full prompt
Adapt the inputs to your task. Remove private data before sending anything to a model. This page copies text locally; it does not run the prompt.
Create an instrumentation plan for [STREAMING WORKFLOW]. Record request start, first visible token, tool start/end, response completion and acceptance completion. Include a cancellation and an interrupted stream. Distinguish observed latency from placeholder values. Specify how partial output is presented so it cannot be mistaken for a verified final result.
Review checklist
- Record first visible token and accepted completion separately.
- Include tool and validation time.
- Test cancellation and interrupted streams.
- Label partial results visibly.
Keep the resource
Download the complete note ↓Markdown · explanation, example, prompt and checklistDownload the prompt ↓Plain text · ready to adaptOpen the full-size visual ↓SVG · scalable reference diagramDocumentation & scope
This is an original workflow template with an illustrative example. It does not report completed model tests or measured performance. Provider APIs can change; check the linked documentation for your exact integration.