Back to evaluations
Public evaluation
Sentiment Consistency Check
Moderation accuracy audit
Evaluation type
task based
Challenge
Automated Content Moderation Workflow with OpenAI Agents SDK
Difficulty
Intermediate
Rigor
Unspecified
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
1
Benchmarks
0
Criteria
0
Task templates
Inputs and expected outputs.
Task 1
Sentiment Consistency Check
Evaluates if the moderation team correctly identifies tone using Hume AI
Input format
Voice recording clip
Output format
Boolean flag and explanation