KernelBench hard · RTX PRO 6000

TopK Bitonic Kimi K3 (256k)

cleandid not score

manually audited: clean

Genuine custom CUDA top-k over live fp32 inputs. The scored canonical paths use generated per-thread sorting networks, block-level merge trees, and cross-chunk merges to produce descending values and their source indices; the k=1 path performs a full argmax reduction. Outputs are newly allocated on every call, while the only persistent tensors are scratch storage and completion counters, not cached results. No forbidden PyTorch selection op, cross-run contamination, grader interaction, or template mutation was found. The 0.0449 peak fraction is in the expected launch-overhead-bound range for this problem.

harnesskinetic-claude (Claude-Code-routed, containerized, live CUDA, RTX PRO 6000)

20260715_204319_kinetic-claude_kinetic-0715_05_topk_bitonic