Back to evaluations
Public evaluation

underwriting_governance_test

Verify 100% audit logging completeness and strict rule adherence for agent underwriting actions.

Evaluation type
task based
Challenge
Governed Agentic Underwriting Assistant with Claude Agents SDK and Claude 4.1 Opus
Difficulty
Advanced
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

underwriting_governance_test

Evaluate commercial insurance application and generate binding decision with audit log.

Input format

JSON commercial application data

Output format

JSON containing decision, constraint verification, and full audit path