Nonprofit AI field kit
Browser-safe methodTen-case evaluation set
A team needs repeatable evidence before a pilot, purchase, workflow change, or model update.
Ten representative cases, expected characteristics, a rubric, and recorded failure patterns.
- 01
Choose normal, difficult, ambiguous, incomplete, multilingual, accessibility, and stop-path cases.
- 02
Use safe synthetic or approved information and document what a good response must contain or avoid.
- 03
Score the current process and proposed workflow with the same outcome-focused rubric.
- 04
Save results, reviewer disagreement, near misses, changes, and the conditions that require retesting.