Back to evaluations
Public evaluation

Threat Classification Test

Ensure agents accurately classify threats and adhere to safety policies.

Evaluation type
task based
Challenge
Develop Secure Multi-Agent Security Operations with LangChain and Claude Opus 4.6
Difficulty
Advanced
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

Threat Classification Test

Evaluate agent capability to flag unauthorized access attempts

Input format

Text log snippet

Output format

Categorization + Risk Score