Back to evaluations
Public evaluation
image_forensics_task
Evaluates forensic image manipulation detection rate, false positive balance, and Mastra state execution integrity.
Evaluation type
task based
Challenge
Image Integrity Verification Agent with Mastra AI and Llama 3.3 70B
Difficulty
Advanced
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
image_forensics_task
Evaluates detection of duplicated or spliced western blot figure regions.
Input format
JSON containing image URL and metadata.
Output format
JSON containing manipulation detected boolean, confidence, and anomaly list.