Back to evaluations
Draft evaluation

Triton Kernel Quantization Benchmark Agent with Pydantic AI and Together AI — evaluation

Evaluates speedup multiplier and model quality preservation for benchmarked kernels.

Evaluation type
task based
Challenge
Triton Kernel Quantization Benchmark Agent with Pydantic AI and Together AI
Difficulty
Advanced
Rigor
Not declared

The author has not specified a rigor level.

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

Kernel Benchmark Task

Runs Triton GEMM kernel benchmarks for matrix shapes M=4096, N=4096, K=4096 under FP16 vs FP8 vs INT4.

Input format

JSON with matrix dimensions and quantization format

Output format

JSON with speedup_factor and perplexity_delta