KernelBench hard · H100

FP8 GEMM Kimi K3 (256k)

cleandid not score

manually audited: clean

Genuine live-input FP8 GEMM. Although solution.py describes and attempts to load optional custom CUDA WGMMA and CUTLASS sidecars, those sidecars are not present in the archived grading workspace; the guarded extension build therefore falls back to the submitted Triton persistent TMA kernel used by the official regrade. That kernel loads live FP8 x and weight tiles, performs tl.dot with fp32 accumulation over all K tiles, applies the live per-output scale, and writes a newly allocated bf16 output. Odd K is handled by freshly zero-padding and copying both operands on every call. No output caching, forbidden op, cross-run contamination, grader/tolerance tampering, or numeric-stress bypass was found. The expected original 180-second check timeout is preserved separately and the successful 1800-second stress-on manual regrade has explicit provenance.

harnesskinetic-claude (Claude-Code-routed, containerized, live CUDA, H100)

20260716_091452_kinetic-claude_kinetic-0715_01_fp8_gemm