KernelBench mega · RTX PRO 6000

rl grid ppo Kimi K3 (1M)

bugdid not score

Ungraded infra casualty, not a hack and not contaminated. The agent session was cut at exactly the 21,600-second harness boundary (harness_exit_code 124, session_complete false), and the official check.py then never executed: the gpu_lock.log ends with "lock_timeout pid=712973 cmd=check.py wait_timeout_s=7200" after a two-hour wait on outputs/gpu.lock, which was held by the stuck sibling run 20260716_150026_..._02_kimi_linear_decode check.py (visible in the later run's ps diagnostics at 4h+ elapsed). result.json correctly records correct=false, failure_reason=timeout, peak_fraction=null, template_mutated=false. This run has NO score; any report attributing 28.7578 to this run id is wrong — that number exists only in the 20260716_233413 sibling's benchmark.log and result.json. The banked solution.py (md5 6a850095479b4ac27e046d4fa82d0ab1, 26,071 bytes) is a genuine in-progress PPO megakernel and is a different file from the sibling run's solution (md5 a63618ba936472df072c942b6e783e69, 32,674 bytes).

harnesskinetic-claude

20260716_150001_kinetic-claude_kinetic-0715_1m__01_rl_grid_ppo