Back to evaluations
Draft evaluation
Image Integrity Verification Agent with Mastra AI and Llama 3.3 70B — evaluation
Evaluates forensic image manipulation detection rate, false positive balance, and Mastra state execution integrity.
Evaluation type
task based
Challenge
Image Integrity Verification Agent with Mastra AI and Llama 3.3 70B
Difficulty
Advanced
Rigor
Not declared
The author has not specified a rigor level.
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
image_forensics_task
Evaluates detection of duplicated or spliced western blot figure regions.
Input format
JSON containing image URL and metadata.
Output format
JSON containing manipulation detected boolean, confidence, and anomaly list.