Back to evaluations
Public evaluation
autogen_harness_task
Evaluates AutoGen state compression and autoscaling recommendations over a simulated 50-turn agent trajectory.
Evaluation type
task based
Challenge
Long-Running Harness Capacity Planner with AutoGen and Kira Integration
Difficulty
Advanced
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
autogen_harness_task
Tests context compression efficiency and memory threshold trigger.
Input format
JSON with session_turns array, current_vram_usage, and memory_cap_mb
Output format
JSON detailing final_token_count, compression_ratio, scaling_signal_emitted, and preserved_facts_count