kernelbench.com

writeups

long-form posts on design choices, rubric leaks, and methodology.

[ latest ]

Hard: Eight Scored Problems, Frontier Models, Two Rubric Leaks

One Blackwell GPU, a focused CUDA problem deck, real coding-agent CLIs as the harness. Frontier models swept; only GPT-5.5 xhigh solved every problem. Two problems leak the rubric: five models all took the same bf16 shortcut on FP8 GEMM, and the only model that implemented Kahan compensated summation scored lowest of the passes.