Back to evaluations
Public evaluation

apt_triage

Tests agent ability to correctly classify APT threat severity and generate valid containment payloads.

Evaluation type
task based
Challenge
Cyber Threat Intelligence Triage Agent using Claude Agents SDK and Pydantic AI
Difficulty
Advanced
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

apt_triage

Analyzes raw network telemetry logs and outputs validated threat containment decisions

Input format

JSON telemetry log with PCAP highlights and netflow data

Output format

Pydantic JSON model with severity, confidence, and firewall blocking parameters