Back to evaluations
Public evaluation
System Health Monitor
Evaluation of agent responsiveness and diagnostic accuracy.
Evaluation type
task based
Challenge
Real-Time AI Operations Controller with Claude Agents SDK
Difficulty
Advanced
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
System Health Monitor
Evaluates agent response time to synthetic anomaly triggers.
Input format
System logs JSON
Output format
JSON decision structure