Back to evaluations
Public evaluation

Report Validation

Evaluates the structural integrity and logic of the investigation reports.

Evaluation type
task based
Challenge
Automated Financial Crime Investigation with Pydantic AI and Llama 3.3 70B
Difficulty
Intermediate
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

Report Validation

Ensures the generated report matches the Pydantic schema

Input format

JSON transaction log

Output format

Validated Pydantic object