Back to evaluations
Public evaluation

System Health Monitor

Evaluation of agent responsiveness and diagnostic accuracy.

Evaluation type
task based
Challenge
Real-Time AI Operations Controller with Claude Agents SDK
Difficulty
Advanced
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

System Health Monitor

Evaluates agent response time to synthetic anomaly triggers.

Input format

System logs JSON

Output format

JSON decision structure