Back to evaluations
Draft evaluation

Post-Hurricane Claims Fraud Detection with Pydantic AI and Qwen 3 — evaluation

Evaluate claim validation accuracy and fraud classification metrics against labeled test claims.

Evaluation type
task based
Challenge
Post-Hurricane Claims Fraud Detection with Pydantic AI and Qwen 3
Difficulty
Intermediate
Rigor
Not declared

The author has not specified a rigor level.

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

claim_fraud_classification

Evaluate claim data and imagery to produce a structured fraud risk verdict.

Input format

JSON containing claim text, damage description, and image URL

Output format

JSON object matching FraudAssessment Pydantic schema