kernelbench.com

KernelBench hard · H100

FP8 GEMM Qwen 3.8 Max

23.6%geomean peak fraction across shapes

manually audited: clean

Genuine custom Hopper fp8 GEMM. The final solution launches Triton tl.dot on float8_e4m3fn operands with fp32 accumulators, applies the live per-output weight scale, and writes bf16. Large aligned shapes use a persistent TMA kernel; K=4127 uses explicit zero-padding copies before the same GEMM; M<=64 uses a transposed split-K GEMM with a zeroed fp32 atomic workspace and a scale/convert epilogue. There is no vendor GEMM call, constant output, benchmark/check discriminator, output memoization, CUDA graph, identity-keyed reuse, or fake-compute path. The publish-grade isolated regrade records peak_fraction=0.2362 and correct=true.

harnessor-fable

20260803_034407_qwen-claude_qwen3.8-max_01_fp8_gemm