Back to evaluations
Public evaluation
Threat Classification Test
Ensure agents accurately classify threats and adhere to safety policies.
Evaluation type
task based
Challenge
Develop Secure Multi-Agent Security Operations with LangChain and Claude Opus 4.6
Difficulty
Advanced
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
Threat Classification Test
Evaluate agent capability to flag unauthorized access attempts
Input format
Text log snippet
Output format
Categorization + Risk Score