Agentic reinforcement fine-tuning uses graded multi-step behavior to change an agent policy.
Versalist prepares evaluation evidence. It does not provide training compute or operate the customer agent runtime.
Identify the skill or policy used by each Episode.
Bind each Episode to one Challenge version.
Test whether graders reward the intended behavior.
Measure one policy revision on a holdout set.
1. Control the comparison
Use these controls before you collect or compare results.
2. Use the procedure
Complete each step in order. Stop when a required input or control is missing.
3. Keep the evidence
Store enough evidence for another reviewer to repeat the decision.
| Record | Required evidence | Failure signal |
|---|---|---|
| Policy source | Skill version and content digest | The training input is unknown |
| Episode source | Challenge, steps, attempt identity | Results mix stale attempts |
| Reward | Grader version and validation cases | The signal rewards a shortcut |
| Rights | Eligibility and provenance record | Private or external data enters the corpus |
| Decision | Baseline, candidate, and holdout result | The revision has no comparison |
4. Keep the product boundary
- Current Trace events contain bounded metadata, hashes, latency, token counts, and sanitized errors.
- Current Trace events do not contain raw prompts, full transcripts, hidden reasoning, or tool payloads.
- An Episode stores rubric scores. It does not store one Episode reward column.
- Versalist does not execute the training job or operate the customer agent runtime.