A Challenge is an evaluation contract. It combines a task, cases, rubric, environment blueprint, and execution policy.
A useful Challenge detects behavior differences and keeps each comparison reproducible.
Write an observable completion condition.
Keep ranking cases private and rights-cleared.
Do not change scoring during a comparison.
Record each material execution condition.
1. Control the comparison
Use these controls before you collect or compare results.
2. Use the procedure
Complete each step in order. Stop when a required input or control is missing.
3. Keep the evidence
Store enough evidence for another reviewer to repeat the decision.
| Record | Required evidence | Failure signal |
|---|---|---|
| Challenge | Versioned task and constraints | The required behavior is vague |
| Cases | Source, rights, version, and visibility | Case provenance is unknown |
| Rubric | Locked criteria and weights | A host can change scores silently |
| Baseline | Complete Episode result | The Challenge has no reference result |
| Maintenance | Discrimination and retirement record | Solved cases stay in ranking |
4. Keep the product boundary
- A Challenge is not a prompt, a leaderboard, or a harness.
- A harness runs the agent. It does not define the evaluation contract.
- Public cases do not provide a trustworthy ranking by themselves.
- Available Trace metadata does not include raw prompts, full transcripts, hidden reasoning, or tool payloads.