Back to evaluations
Public evaluation
claim_fraud_classification
Evaluate claim validation accuracy and fraud classification metrics against labeled test claims.
Evaluation type
task based
Challenge
Post-Hurricane Claims Fraud Detection with Pydantic AI and Qwen 3
Difficulty
Intermediate
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
claim_fraud_classification
Evaluate claim data and imagery to produce a structured fraud risk verdict.
Input format
JSON containing claim text, damage description, and image URL
Output format
JSON object matching FraudAssessment Pydantic schema