Back to evaluations
Public evaluation

Reasoning Chain Audit

Multi-step reasoning accuracy and retrieval performance.

Evaluation type
task based
Challenge
Extended Reasoning Orchestration with Claude Agents SDK and Milvus
Difficulty
Advanced
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

Reasoning Chain Audit

Validates the agent output matches the ground truth logic for long documents

Input format

Complex document text

Output format

Structured synthesis report