# AI model releases: turn an announcement into a test

Separate the provider claim from the application question it raises. Preserve the release date and model version, then define one bounded trial on your workload. Do not turn a vendor comparison into your own measured result.

## Worked example

A release claims lower task cost. A useful trial measures all attempts, tool usage and acceptance on your own fixed set. Until those runs exist, the worksheet should show a proposed test and unknown results.

## Copyable prompt

```text
Review [OFFICIAL RELEASE] for [WORKLOAD]. Extract dated provider claims with source URLs. Convert each relevant claim into a testable application question. Define a fixed dataset, baseline, acceptance rule, run count and measurements. Keep result cells blank until measured. Flag claims that the available setup cannot test fairly.
```

## Checklist

- Record release date, source and exact model.
- Keep vendor claims separate from observations.
- Use the same task and acceptance criteria.
- Publish the method and limits alongside any result.

## Further reading

[OpenAI: evaluation best practices](https://developers.openai.com/api/docs/guides/evaluation-best-practices)
