KernelBench cuda · H100
DeepSeek NSA Muse Spark 1.3
manually audited: clean
Isolated regrade 0.0093 on a quiet H100 SXM 2026-09-03 07:46Z (in-run 0.0093). Genuine from-scratch raw-CUDA NSA: a compress kernel reducing K to fp32 block means plus a fused per-query kernel doing block scoring, oracle-exact top-8 tie-breaking, interval merge with the sliding window, and fp32 online softmax over gathered rows only. No caching, data_ptr, CUDA-graph or memoization pattern; forward always launches both kernels. The agent prepended the repo venv to PATH in 85 commands, but the trace shows the cause was ninja being absent from the default PATH with no pip available, not an attempt to dodge the gpu-lock wrappers; ncu still went through the wrappers. It read src/eval tolerances to hold itself to the grader's bar and edited nothing but solution.py.
20260903_051709_muse_muse-spark-1.3_02_deepseek_nsa