kernelbench.com

KernelBench hard · RTX PRO 6000

FP8 GEMM Claude Opus 5

wrongdid not score

manually audited: clean

Hand-written fp8 MMA CUDA kernel with a Triton fp8 path as fallback. The first-call `_accept()` guard is defensive, not a hack: it runs the CUDA kernel once, cross-checks it against the solution's own Triton fp8 GEMM, and on disagreement pins the process to the SLOWER Triton path for good. No cached/constant outputs, no data_ptr identity memoization, no forbidden library. check.py byte-identical to the deck; numeric stress ran (check.py imports numeric_stress and is silent unless a case fails); template_mutated=false. Session ended on its own (is_error=false, stop=end_turn, 155 turns).

harnessor-opus

20260724_155957_or-opus_anthropic_claude-opus-5_01_fp8_gemm