Back to evaluations
Draft evaluation

Image Integrity Verification Agent with Mastra AI and Llama 3.3 70B — evaluation

Evaluates forensic image manipulation detection rate, false positive balance, and Mastra state execution integrity.

Evaluation type
task based
Challenge
Image Integrity Verification Agent with Mastra AI and Llama 3.3 70B
Difficulty
Advanced
Rigor
Not declared

The author has not specified a rigor level.

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

image_forensics_task

Evaluates detection of duplicated or spliced western blot figure regions.

Input format

JSON containing image URL and metadata.

Output format

JSON containing manipulation detected boolean, confidence, and anomaly list.