Back to evaluations
Public evaluation

Benchmark-aligned evaluation

Factual Accuracy and Grounding

Evaluation type
benchmark aligned
Challenge
Industrial Automation: Develop a system for factory robots to adapt to new manufacturing tasks
Difficulty
Advanced
Rigor
Production-grade with specific benchmark targets

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
0
Benchmarks
4
Criteria
2

Evaluation method

Approach

Official AI benchmark-aligned evaluation

Criteria

What gets scored.

Criterion 1

Factual Accuracy and Grounding

Evaluation of factual correctness and information grounding

Metrics
  • SimpleQA Accuracy: Direct factual question answering
  • FACTS Grounding Score: Citation and source attribution accuracy
  • Hallucination Rate: Percentage of unsupported claims
  • Source Verification: Accuracy of cited references
  • Temporal Accuracy: Correctness of time-sensitive information
Benchmarks
  • SimpleQA
  • FACTS Grounding
Target

SimpleQA > 85%, Hallucination Rate < 5%, FACTS Grounding > 80%

Evaluation type

quantitative

Criterion 2

Robustness and Reliability

System stability, consistency, and adversarial resistance

Metrics
  • Output Consistency: Variance across similar inputs
  • Adversarial Robustness: Resistance to malicious inputs
  • Performance Stability: Consistency across time and load
  • Failure Graceful Degradation: Behavior under stress
  • Safety Compliance: Adherence to safety guidelines
Benchmarks
  • RobustBench
  • SafetyBench
Target

Consistency > 95%, Adversarial Success < 5%

Evaluation type

quantitative