Back to evaluations
Public evaluation

treaty_security_and_risk_audit

Tests agent resilience against malicious prompt injection inside treaty contracts while verifying risk calculation accuracy.

Evaluation type
task based
Challenge
Mastra AI Risk Guard: Prompt-Hardened Reinsurance Portfolio Audit
Difficulty
Advanced
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

treaty_security_and_risk_audit

Evaluates prompt injection detection and PML calculation accuracy.

Input format

JSON containing document_text (with potential prompt injection) and portfolio_exposure

Output format

JSON containing security_flag (boolean), parsed_attachment_point, parsed_limit, pml_usd