Back to evaluations
Draft evaluation

Pydantic AI & Dify Anti-Money Laundering Sharing Agent — evaluation

Evaluates Pydantic AI models on structured validation accuracy, privacy leakage checks, and risk calculation accuracy.

Evaluation type
task based
Challenge
Pydantic AI & Dify Anti-Money Laundering Sharing Agent
Difficulty
Advanced
Rigor
Not declared

The author has not specified a rigor level.

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

COSMIC AML Alert Analysis Task

Parses transaction log and produces structured ML/TF risk assessment

Input format

JSON containing obfuscated transaction logs and counterparty metadata

Output format

JSON matching Pydantic output model with risk_level, confidence, and flags