Back to evaluations
Public evaluation
Reasoning Chain Audit
Multi-step reasoning accuracy and retrieval performance.
Evaluation type
task based
Challenge
Extended Reasoning Orchestration with Claude Agents SDK and Milvus
Difficulty
Advanced
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
Reasoning Chain Audit
Validates the agent output matches the ground truth logic for long documents
Input format
Complex document text
Output format
Structured synthesis report