KernelBench mega · H100
rl grid ppo Kimi K3 (1M)
manually audited: clean
Clean and megakernel-authentic at 3.6765x (91,911,815 SPS vs the 25,000,000 SPS anchor) on H100. The entire PPO training run — all iterations of the 32-step x 4096-env rollout with fused GAE, advantage normalization, and 4 epochs x 4 minibatches of exact PPO gradient + global grad-norm clip + Adam — executes inside ONE cudaLaunchCooperativeKernel launch per train() call (extension ppo_megakernel_v18); launch count does not scale with env steps. No CUDA graphs, no caching keyed on inputs, no torch.compile, no RL library, no check.py sniffing, and all five grader templates byte-match the harness snapshot. This was a 6h-capped session (harness_exit_code=124, session_complete=false): the agent was cut off mid-optimization, which is why this cell sits well below the later 23.1x sibling — the kernel graded is simply the best checkpoint the agent had written before the wall.
20260716_150032_kinetic-claude_kinetic-0715_1m__01_rl_grid_ppo