Back to evaluations
Public evaluation
Benchmark-aligned evaluation
Factual Accuracy and Grounding
Evaluation type
benchmark aligned
Challenge
Industrial Automation: Develop a system for factory robots to adapt to new manufacturing tasks
Difficulty
Advanced
Rigor
Production-grade with specific benchmark targets
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Tasks
0
Benchmarks
4
Criteria
2
Evaluation method
Approach
Official AI benchmark-aligned evaluation
Criteria
What gets scored.
Criterion 1
Factual Accuracy and Grounding
Evaluation of factual correctness and information grounding
Metrics
- SimpleQA Accuracy: Direct factual question answering
- FACTS Grounding Score: Citation and source attribution accuracy
- Hallucination Rate: Percentage of unsupported claims
- Source Verification: Accuracy of cited references
- Temporal Accuracy: Correctness of time-sensitive information
Benchmarks
- SimpleQA
- FACTS Grounding
Target
SimpleQA > 85%, Hallucination Rate < 5%, FACTS Grounding > 80%
Evaluation type
quantitative
Criterion 2
Robustness and Reliability
System stability, consistency, and adversarial resistance
Metrics
- Output Consistency: Variance across similar inputs
- Adversarial Robustness: Resistance to malicious inputs
- Performance Stability: Consistency across time and load
- Failure Graceful Degradation: Behavior under stress
- Safety Compliance: Adherence to safety guidelines
Benchmarks
- RobustBench
- SafetyBench
Target
Consistency > 95%, Adversarial Success < 5%
Evaluation type
quantitative