KernelBench hard · RTX PRO 6000
Paged Attention Kimi K3 (256k)
manually audited: clean
Genuine hand-written GQA paged-attention decode in CUDA C++ via load_inline. The fast path launches one CTA per batch, KV-head group, sequence split, and optional query-head split; follows the live randomized block_table for every KV page; streams packed K and V; computes QK dots and fp32 online softmax in the exp2 domain; accumulates weighted V; and merges token-stream and sequence-split partials. A generic CUDA fallback performs the same live paged gather and online softmax for unsupported dimensions. Output is newly allocated for every call. Persistent tensors are only self-resetting split-reduction scratch and counters, not cached results. No forbidden library or dense-attention shortcut is used. The 0.4489 geomean peak fraction is physically plausible and exactly recomputes from the five logged timings.
20260715_204254_kinetic-claude_kinetic-0715_03_paged_attention