Back to evaluations
Draft evaluation

Automated Image Misconduct Detection RAG Pipeline with LlamaIndex — evaluation

Evaluates candidate capability to index figures, detect manipulated panels, and output valid structured audit summaries.

Evaluation type
task based
Challenge
Automated Image Misconduct Detection RAG Pipeline with LlamaIndex
Difficulty
Advanced
Rigor
Not declared

The author has not specified a rigor level.

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

Image Audit Verification

Evaluates detection accuracy for spliced blot images and JSON audit output format.

Input format

JSON containing figure manifest with image file paths

Output format

JSON containing detected duplicate regions, similarity scores, and flag boolean