Back to evaluations
Public evaluation
purdue_audit_eval
Evaluates CrewAI team output for accuracy in detecting illegal cross-level communications.
Evaluation type
task based
Challenge
Purdue Model Multi-Agent Enterprise OT/IT Architecture Team with CrewAI
Difficulty
Intermediate
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
purdue_audit_eval
Audits netflow logs between OT levels and outputs compliance violations.
Input format
JSON array of network flow tuples (src_ip, src_level, dst_ip, dst_level, protocol)
Output format
JSON report detailing severe_violations, recommended_dmz_rules, and risk_score