kernelbench.com

KernelBench hard · RTX PRO 6000

Paged Attention Claude Opus 5

wrongdid not score

manually audited: clean

Real CUDA paged-attention decode kernel: cp.async shared-memory staging, one query head per warp, online softmax over pages. No F.softmax / scaled_dot_product_attention / flash-attn fallback, no cached outputs. check.py byte-identical to the deck; numeric stress ran; template_mutated=false. 0.68 of peak is high but plausible for a well-tuned bandwidth-bound decode kernel. Session ended on its own (is_error=false, stop=end_turn, 445 turns).

harnessor-opus

20260724_160047_or-opus_anthropic_claude-opus-5_03_paged_attention