Industrial Smelter Restoration Monitor with Claude Agents SDK
Emirates Global Aluminium's Al-Taweelah restoration project requires automated tracking of repair progress across heavy industrial assets. Build an intelligent visual and diagnostic inspection agent using Claude Agents SDK and CleverBee to inspect repair logs and sensor feeds. Deliver precise completion verification reports with 90% diagnostic accuracy.
What you are building
The core problem, expected build, and operating context for this challenge.
Develop an inspection analysis pipeline leveraging Claude Agents SDK with extended thinking to verify industrial smelter restoration milestones.
How work is evaluated
Measures reasoning accuracy and completion rate calculations against ground-truth site audit reports.
Shared data for this challenge
Review public datasets and any private uploads tied to your build.
How submissions are scored
These dimensions define what the evaluator checks and which criteria separate a passable run from a strong one.
extended_thinking_verification
Checks that agent response includes structured reasoning trace steps
This dimension contributes its full weight only when the submission satisfies the requirement. Partial credit is not awarded.
completion_estimation_accuracy
Accuracy of physical completion calculation vs master auditor score • target: 0.9 • range: 0-1
This dimension contributes its full weight only when the submission satisfies the requirement. Partial credit is not awarded.
What you should walk away with
Utilize Claude Agents SDK with extended thinking for multi-step reasoning
Integrate CleverBee framework tools for multi-agent system execution
Parse visual inspection logs and operational telemetry from smelting potlines
Validate physical repair completion percentage against project baseline documentation
Reference links and supporting material
Dataset containing potline reconstruction telemetry, non-destructive testing reports, and thermal imaging annotations for 100 smelter cells.
How this agent runs
Measures reasoning accuracy and completion rate calculations against ground-truth site audit reports.
Challenge input
JSON containing inspection summaries, thermal sensor logs, and image features
Claude Agents SDK
Offers deep reasoning via extended thinking and multimodal understanding.
CleverBee
Facilitates modular agent node management and streaming execution.
Evaluated output
JSON with verified progress percentage, anomaly count, and sign-off status
- Checks that agent response includes structured reasoning trace steps
- Accuracy of physical completion calculation vs master auditor score • target: 0.9 • range: 0-1
- Benchmark: IndustrialInspectBench
- Completion Estimation Accuracy target: 0.9
- 1 public reference case
- Python execution harness
- Python sandbox (unavailable on Versalist)
View technical recipe
Configured tools
- CleverBee · Required
- Agno · Optional
- BoTorch · Optional
- Agno · Optional
Evaluation contract
- extended_thinking_verification · Weight 1
- completion_estimation_accuracy · Weight 1
Recipe state
This is a preview. The configuration can change before the evaluation recipe is locked.