Evaluation Data Quality

Build rights-cleared evaluation cases with stable labels, provenance, and review rules.

Read time
4 min
Scope
0 sections
Level
Intermediate / 1 of 2 in track

Outcome: Create an evaluation dataset that supports repeatable comparisons.

Evaluation quality depends on the cases, labels, source rights, and review history.

A large dataset is not useful when the task contract or provenance is weak.

Decision
Can another reviewer explain where each case came from and why its label is correct?
Keep external, self-reported, synthetic, and platform-verified records in separate classes.
Source
Record origin

Keep the source identity and capture time.

Rights
Confirm allowed use

Store the basis for evaluation use.

Label
Use one rubric

Document reviewer agreement and corrections.

Version
Preserve history

Do not rewrite prior evaluation results.

1. Control the comparison

Use these controls before you collect or compare results.

Contract
Start from the evaluation contract
Collect only cases that measure the named behavior and its known risks.
Define case eligibility
Name required metadata
Reject unrelated volume
Evidence
Keep provenance complete
Each case needs a source, capture method, rights class, and transformation record.
Use stable source identifiers
Record transformations
Keep synthetic markers
Review
Control label quality
A label is evidence only when reviewers use the same decision rule.
Write a label rubric
Measure disagreement
Keep adjudication reasons
Decision
Protect the holdout
Separate optimization examples from cases used for final release decisions.
Restrict access
Check duplicate content
Monitor contamination

2. Use the procedure

Complete each step in order. Stop when a required input or control is missing.

1
Define case eligibility
State the behavior, source classes, and exclusion rules.
2
Capture provenance
Record the source, rights, time, and transformation.
3
Write the label rubric
Define each label and its decision boundary.
4
Review a sample
Measure agreement before you label the complete set.
5
Create the split
Separate iteration cases from protected holdout cases.
6
Version the manifest
Record additions, corrections, exclusions, and reviewer decisions.

3. Keep the evidence

Store enough evidence for another reviewer to repeat the decision.

RecordRequired evidenceFailure signal
SourceStable identifier and capture recordThe origin cannot be verified
RightsUse class and restriction recordEvaluation use is not authorized
TransformationCode or recorded operationThe source cannot be reconstructed
LabelRubric, reviewer, and adjudicationLabel meaning changed silently
SplitMembership and access controlHoldout cases entered optimization

4. Keep the product boundary

  • Do not treat synthetic data as observed user behavior.
  • Do not mix unknown-provenance records into the rights-cleared corpus.
  • Do not use public benchmark data as a private holdout.
  • Do not delete corrected labels. Preserve the prior version and reason.

5. Apply this on Versalist

Attach cases to a Challenge
Keep each evaluation case connected to its task and rubric.
Read Challenge Design
Apply the evidence policy
Review the current evidence rights and access rules.
Read result documentation
Release gate
Make the decision explicit
Use the dataset only after provenance, rights, labels, and holdout controls pass review.
More guides

Keep going

Current track: Change discipline