Guides for verified agent changes
Define the task, keep the evaluation contract stable, inspect the evidence, and record the release decision. Each guide supports part of the closed agent-improvement loop.
Start here when a score cannot show whether an agent change is safe.
Use this track before you publish a new instruction, dataset, or policy revision.
Evaluation systems guides
Define tasks, run agents, grade results, inspect failures, and control changes.
Teams that need reproducible evidence before an agent change reaches production.
Start here when a score cannot show whether an agent change is safe.
Start with this guideAgent Evaluation
Build an evaluation contract that links tasks, graders, run conditions, and release decisions.
Teams replacing informal output review with repeatable evaluation evidence.
Create a repeatable release gate for one agent behavior.
Challenge Design
Turn a real agent task into a versioned Challenge with public and hidden cases.
Operators who create public or private agent evaluations.
Write a Challenge specification that can detect useful agent differences.
Agentic Reinforcement Fine-Tuning
Prepare trajectory evidence for reinforcement fine-tuning without overstating current trace coverage.
Teams that can reproduce and grade multi-step agent behavior.
Decide whether your evaluation data is ready for an Agentic RFT experiment.
Change discipline guides
Treat prompts, data, and policy instructions as versioned agent components.
Teams that improve agents and need proof against regressions.
Use this track before you publish a new instruction, dataset, or policy revision.
Start with this guideEvaluation Data Quality
Build rights-cleared evaluation cases with stable labels, provenance, and review rules.
Teams whose evaluation results depend on inconsistent cases or labels.
Create an evaluation dataset that supports repeatable comparisons.
Prompt and Policy Changes
Version prompts and agent instructions with evaluation evidence and rollback criteria.
Teams that change prompts, system instructions, or skill bundles.
Publish one measured instruction change with a clear rollback condition.
Keep going
Build an evaluation contract that links tasks, graders, run conditions, and release decisions.
Turn a real agent task into a versioned Challenge with public and hidden cases.
Prepare trajectory evidence for reinforcement fine-tuning without overstating current trace coverage.
Build rights-cleared evaluation cases with stable labels, provenance, and review rules.