KernelBench cuda · H100
Grid + MinGRU SPS Qwen 3.8 Max
This is a genuine seed-, input-, state-, horizon-, and weight-dependent CUDA implementation, not a constant answer or skipped operation. The official H100 check is unequivocally incorrect, so the cell is an ordinary candidate-code bug rather than an audit rejection. result.json records correct=false, failure_reason=check_failed, check_exit_code=1, and no benchmark execution or peak fraction. The checker first fails at seed 42 run() positions after the policy_forward stress cases and env_step passed. Static review identifies the cause: after a one-way grid barrier, a fast block can start the next step and reset the parity slot before a slower block reads the current any-hit value. Blocks then disagree on globally coupled LCG advancement and eventually diverge in positions. If that synchronization defect is fixed, the identity/version-keyed split-weight cache still requires the isolated same-buffer-overwrite recomputation test specified below.
20260803_194401_or-fable_qwen_qwen3.8-max_04_grid_mingru_sps