Back to evaluations
Public evaluation
Report Validation
Evaluates the structural integrity and logic of the investigation reports.
Evaluation type
task based
Challenge
Automated Financial Crime Investigation with Pydantic AI and Llama 3.3 70B
Difficulty
Intermediate
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
Report Validation
Ensures the generated report matches the Pydantic schema
Input format
JSON transaction log
Output format
Validated Pydantic object