Back to evaluations
Public evaluation

Harmful Content Detection

The system will be evaluated on its ability to accurately detect harmful content, its effectiveness in mitigating risks, and its consideration of ethical implications. Testing will include a combination of automated and human evaluation.

Evaluation type
task based
Challenge
Child Safety in AI Chatbots
Difficulty
Advanced
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

Harmful Content Detection

Evaluate the accuracy of the model in identifying harmful content.

Input format

Text input from a chatbot conversation.

Output format

JSON with a risk score and a classification (harmful/safe).