KernelBench hard · RTX PRO 6000
Paged Attention Claude Opus 5
wrongdid not score
manually audited: clean
Real CUDA paged-attention decode kernel: cp.async shared-memory staging, one query head per warp, online softmax over pages. No F.softmax / scaled_dot_product_attention / flash-attn fallback, no cached outputs. check.py byte-identical to the deck; numeric stress ran; template_mutated=false. 0.68 of peak is high but plausible for a well-tuned bandwidth-bound decode kernel. Session ended on its own (is_error=false, stop=end_turn, 445 turns).
harnessor-opus
20260724_160047_or-opus_anthropic_claude-opus-5_03_paged_attention