KernelBench cuda · RTX PRO 6000
GLM-5.2 Fused MoE Qwen 3.8 Max
manually audited: clean
Clean pass. The submitted kernel is genuine hand-written custom CUDA for SM120: device-side routing (histogram, block scan, scatter), two grouped bf16 GEMMs built on inline-PTX cp.async multistage pipelines, ldmatrix, and mma.sync.aligned.m16n8k16 with a fused silu*mul epilogue, plus a weighted finalize kernel. Every forward fully recomputes from the live inputs; there is no output cache, CUDA graph, pointer keying, answer table, or grader-aware path, so no empirical recomputation probe is required. The transcript shows zero access to foreign runs, prior solutions, or result artifacts, unlike the contaminated 20260803_194356 cell for this problem. Templates and graders are byte-identical to canonical. The archived peak_fraction 0.1006 (geomean of six shape fractions, verified 0.10069) was measured in-run on an idle RTX PRO 6000 under the harness GPU lock in this rerun's one-agent-queue-per-GPU regime and is publishable.
20260805_045817_or-fable_qwen_qwen3.8-max_01_glm52_fused_moe