Back to evaluations
Public evaluation

MCED Guideline Query Task

Evaluates the factual fidelity and clinical precision of RAG generated diagnostic follow-up recommendations.

Evaluation type
task based
Challenge
LlamaIndex MCED Diagnostic Utility RAG with GPT-5 Pro
Difficulty
Intermediate
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

MCED Guideline Query Task

Query the RAG index with an MCED signal report and check compliance of recommended workup.

Input format

JSON with patient_age, mced_signal_tissue_of_origin, biomarker_level

Output format

JSON containing recommended_imaging, specialist_referral, and citations