# Tool calling: test the complete round trip

Treat tool selection, execution, result handoff and the final answer as separate checkpoints. A model requesting a function does not prove that the function ran. Keep a correlation identifier so each returned result belongs to the right call.

## Worked example

Example test: request the status of order DEMO-42. A mock tool returns status=delayed. The final answer must report delayed. Repeat with a tool error and verify that the assistant reports the limitation instead of inventing a delivery date.

## Copyable prompt

```text
Write a local test plan for [TOOL INTEGRATION]. Trace the user request, tool-call identifier, validated arguments, tool result and final answer. Use a mock tool and no external writes. Include a success, malformed arguments, a tool timeout, an unknown tool and an out-of-order result. State the expected behavior for each case. Check the current provider API documentation before proposing SDK code.
```

## Checklist

- Validate arguments before calling the tool.
- Match the result to its call identifier.
- Exercise tool failures as well as success.
- Check the final answer against the actual result.

## Further reading

[OpenAI: function calling](https://developers.openai.com/api/docs/guides/function-calling)
