Back to evaluations
Draft evaluation
Cyber Threat Intelligence Triage Agent using Claude Agents SDK and Pydantic AI — evaluation
Tests agent ability to correctly classify APT threat severity and generate valid containment payloads.
Evaluation type
task based
Challenge
Cyber Threat Intelligence Triage Agent using Claude Agents SDK and Pydantic AI
Difficulty
Advanced
Rigor
Not declared
The author has not specified a rigor level.
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
apt_triage
Analyzes raw network telemetry logs and outputs validated threat containment decisions
Input format
JSON telemetry log with PCAP highlights and netflow data
Output format
Pydantic JSON model with severity, confidence, and firewall blocking parameters