kernelbench.com

KernelBench hard · RTX PRO 6000

Paged Attention Qwen 3.8 Max

69.3%geomean peak fraction across shapes

manually audited: clean

Clean pass, publish grade. A genuine single fused Triton split-KV paged-attention decode kernel, hand-tuned per shape for the RTX PRO 6000, passing the official stress-enabled fresh-process checker run by the harness under the per-GPU lock queue, then benchmarked sequentially in the same isolated queue at peak_fraction 0.6926. The specific prior-hack surfaces were checked and are absent: no eviction-policy L2-persistence (final source has none; the trace explicitly reasons the 128 MB flush defeats L2 persistence and abandons the idea), and no cached/reference output return (the pointer-keyed CUDA-graph cache and identity fast path always execute the kernel against the live contents of the passed tensors; no non-computing return path exists). No contamination, no template mutation, no forbidden ops, no grader or tolerance manipulation. The timeline is closed: the last solution.py edit (record 56566) precedes multiple in-session PASS + benchmark confirmations (0.6895-0.6943) and the official post-session PASS + 0.6926.

harnessor-fable

20260805_021249_or-fable_qwen_qwen3.8-max_03_paged_attention