Back to evaluations
Draft evaluation

Long-Running Harness Capacity Planner with AutoGen and Kira Integration — evaluation

Evaluates AutoGen state compression and autoscaling recommendations over a simulated 50-turn agent trajectory.

Evaluation type
task based
Challenge
Long-Running Harness Capacity Planner with AutoGen and Kira Integration
Difficulty
Advanced
Rigor
Not declared

The author has not specified a rigor level.

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

autogen_harness_task

Tests context compression efficiency and memory threshold trigger.

Input format

JSON with session_turns array, current_vram_usage, and memory_cap_mb

Output format

JSON detailing final_token_count, compression_ratio, scaling_signal_emitted, and preserved_facts_count