KernelBench cuda · H100
DeepSeek NSA Qwen 3.8 Max
manually audited: clean
Clean cell. The passing 0.0452 implementation is genuine custom CUDA written entirely inside this sandboxed rerun: a six-kernel pipeline (block-mean/ prefix-sum prep, warp top-8 select with the reference tie-break, bucket scan/fill, a tensor-core m16n8k16 scatter attention kernel with per-row spinlocked online-softmax merge, and a finalize divide) developed and debugged in-session over 4h18m through three architecture generations (v1 scatter -> v2 -> MMA rewrite). The container workspace contained only this problem plus the shared src tree; the transcript has zero references to any foreign run id, outputs/runs path, sibling problem, or network fetch, so the contamination that invalidated the 20260803 Qwen NSA cells is absent here. All seven grader files and the entire repo/src tree are byte-identical to the canonical H100 deck. Forward recomputes everything from live q/k/v on every call with a fresh output tensor: no pointer-keyed cache, CUDA graph, or input-identity behavior exists, so no empirical recompute is required. Grading ran on the dedicated per-GPU rerun queue with zero lock wait immediately after session end; publish_grade is true.
20260805_193333_or-fable_qwen_qwen3.8-max_02_deepseek_nsa