Back to evaluations
Public evaluation
underwriting_governance_test
Verify 100% audit logging completeness and strict rule adherence for agent underwriting actions.
Evaluation type
task based
Challenge
Governed Agentic Underwriting Assistant with Claude Agents SDK and Claude 4.1 Opus
Difficulty
Advanced
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
underwriting_governance_test
Evaluate commercial insurance application and generate binding decision with audit log.
Input format
JSON commercial application data
Output format
JSON containing decision, constraint verification, and full audit path