KernelBench hard · RTX PRO 6000

Sonic MoE Kimi K3 (256k)

cleandid not score

manually audited: clean

Genuine grouped GEMM plus fused SwiGLU implemented in CUTLASS CuteDSL, which is a legitimate custom-kernel path on KernelBench-Hard. A persistent SM120 kernel builds its expert-tile schedule from live expert_offsets, TMA-loads live hidden-state and per-expert gate/up weights through a three-stage pipeline, performs both bf16 GEMMs with fp32 mma.sync accumulation, applies silu(gate)*up in registers, and predicated-stores a freshly allocated bf16 output. The only cache is the compiled kernel executable; its callable is invoked unconditionally with current pointers on every forward. No forbidden PyTorch/vendor op, cached result, cross-run artifact, grader modification, or numeric-stress bypass is present. The recorded 0.0659 is plausible and exactly reproduced by the three logged per-shape fractions. The initial failed grade was an environment-only missing-CUTLASS dependency after the grading venv was rebuilt; the archived solution was unchanged for the successful numeric-stress regrade.

harnesskinetic-claude (Claude-Code-routed, live CUDA, RTX PRO 6000)

20260716_112653_kinetic-claude_kinetic-0715_06_sonic_moe_swiglu