KernelBench hard · H100

Sonic MoE Kimi K3 (256k)

cleandid not score

manually audited: clean

Genuine MoE up-projection: a hand-rolled SM90 (Hopper) warp-specialized grouped GEMM in inline CUDA/PTX — TMA (cp.async.bulk.tensor) + mbarrier 4-stage pipeline, 1 producer + 2 consumer warpgroups issuing WGMMA m64n128k16 bf16->f32 into dual gate/up accumulators, cluster-2 TMA multicast of weight tiles, a device scheduler kernel expanding live expert_offsets into swizzled per-tile work, and a silu(g)*u epilogue straight from registers — plus a fully-masked Triton fallback for off-alignment shapes. No sonic_moe import, no torch.matmul/bmm/F.linear, no cached or constant output path, no grader interaction. The 0.0784 fraction is a real timing. Result was manually regraded after the original grading process died (orphaned wrapper); regrade ran with numeric stress on and passed.

harnesskinetic-claude (Claude-Code-routed, containerized, live CUDA)

20260715_204056_kinetic-claude_kinetic-0715_06_sonic_moe_swiglu