Back to evaluations
Public evaluation
upi_fraud_triage_eval
Evaluates the accuracy and latency of the UPI fraud agent using simulated payment event streams.
Evaluation type
task based
Challenge
UPI Fraud Detection Agent with OpenAI Agents SDK and Langfuse
Difficulty
Advanced
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
upi_fraud_triage_eval
Evaluates transaction payloads and assigns a risk label with confidence score.
Input format
JSON transaction object containing amount, payer_id, payee_id, location, and device_id.
Output format
JSON object with risk_status ('APPROVE', 'FLAG', 'BLOCK'), risk_score (0-100), and rationale string.