Back to evaluations
Public evaluation

sanctions_filter_eval

Evaluates stablecoin settlement approval logic and compliance adherence under synthetic high-volume transaction stress.

Evaluation type
task based
Challenge
Stablecoin Settlement Guardrails with OpenAI Agents SDK and GPT-5 Pro
Difficulty
Advanced
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

sanctions_filter_eval

Evaluates agent ability to detect flagged wallet addresses and illegal entity combinations.

Input format

JSON containing sender, receiver, amount, and asset details

Output format

JSON with decision ('APPROVE'|'REJECT'), reason, and confidence score