KernelBench cuda · RTX PRO 6000
GLM-5.2 Fused MoE Kimi K3 (256k)
cleandid not score
manually audited: clean
Genuine hand-written SM120 CUDA/PTX fused MoE: device-side expert histogram/scan/scatter, paired gate/up bf16 MMA with fused SwiGLU, down-projection MMA, and weighted token reduction. All kernels consume live activations, routing, expert weights, and model parameters on every call; outputs and workspaces are freshly allocated. No Triton/DSL, forbidden op, output cache, grader tampering, or foreign-kernel contamination. The 0.0446 geomean is a real solution timing.
harnesskinetic-claude (Claude-Code-routed, live CUDA, RTX PRO 6000)
20260716_112833_kinetic-claude_kinetic-0715_01_glm52_fused_moe