Design an evaluation for [TASK] using [INPUT CASES]. Define an explicit pass/fail rule for each case before running anything. Record model version, prompt version, tool access, attempt count, total cost and time to an accepted result. Include all failed attempts. Separate measured values from unknown values. Return a blank scorecard and a repeatable run protocol. Do not choose a winner without completed trials.