Back to evaluations
Public evaluation

quantization_benchmark_task

Evaluates kernel benchmark reporting accuracy and cost trade-off calculation fidelity.

Evaluation type
task based
Challenge
AI SDK & OpenAI o3 Quantization Kernel Benchmarking Suite
Difficulty
Intermediate
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

quantization_benchmark_task

Benchmarks GEMM execution time for FP16 vs FP8 vs INT4 across varying matrix dimensions.

Input format

JSON containing matrix_dimensions, precision_formats, batch_sizes

Output format

JSON containing latency_matrix_ms, memory_footprint_mb, perplexity_delta, recommended_format