# AI evaluation cases: turn a failure into a regression test

Preserve the input, observed failure and expected behavior before editing the prompt. Add a neighboring case that should still succeed so the fix is not narrowly fitted to one example. Keep the original failure in future regression runs.

## Worked example

Synthetic failure: an assistant returns a confident policy answer when the source is absent. A regression case should expect an explicit limitation. A nearby control case supplies the policy and expects a supported answer.

## Copyable prompt

```text
Convert [FAILURE REPORT] into an evaluation case. Record input, relevant context, observed output, expected behavior and a grading rule. Add one nearby case that should still pass and one ambiguous case that should be escalated. Do not fabricate a measured success rate. Explain which assertion can be automated and which requires review.
```

## Checklist

- Preserve the failure before changing the prompt.
- Write expected behavior in observable terms.
- Add a nearby successful control case.
- Keep holdout cases outside prompt tuning.

## Further reading

[OpenAI: evaluation best practices](https://developers.openai.com/api/docs/guides/evaluation-best-practices)
