Back to evaluations
Draft evaluation
Automated Image Misconduct Detection RAG Pipeline with LlamaIndex — evaluation
Evaluates candidate capability to index figures, detect manipulated panels, and output valid structured audit summaries.
Evaluation type
task based
Challenge
Automated Image Misconduct Detection RAG Pipeline with LlamaIndex
Difficulty
Advanced
Rigor
Not declared
The author has not specified a rigor level.
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
Image Audit Verification
Evaluates detection accuracy for spliced blot images and JSON audit output format.
Input format
JSON containing figure manifest with image file paths
Output format
JSON containing detected duplicate regions, similarity scores, and flag boolean