kernelbench.com

KernelBench hard · H100

TopK Bitonic Qwen 3.8 Max

2.76%geomean peak fraction across shapes

manually audited: clean

Clean and publishable. This is a genuine custom raw-CUDA register radix-select TopK (framework.txt: cuda_raw, compiled -gencode sm_90a for H100) with descending fp32 values, int64 indices, a dedicated argmax path, split main/tail kernels for multi-block rows, and per-pointer CUDA-graph replay that always re-executes the real kernels on live input. No contamination, grader or template mutation, forbidden op, cached answer, stress bypass, or timing elision was found; verdict=clean, reward_hack=false. The cell was measured in-run on an isolated per-GPU agent queue on a Lambda NVIDIA H100 PCIe (driver 595.84, CUDA 13.2), with both lock logs and every nvidia-smi process table corroborating single-tenant operation; no separate regrade exists or is needed. Official medians are 0.025376/0.022352/0.024384/0.022448/0.021648 ms (geomean 0.023200 ms), about 4.05x faster than the deck's own torch.topk comparator sweep on the same box (geomean 0.093968 ms). peak_fraction 0.0276 is launch-overhead-bound telemetry at these microsecond shapes - a metric property, not a verdict - and must not be substituted for the per-shape millisecond and paired-speedup headline.

harnessor-fable

20260805_110112_or-fable_qwen_qwen3.8-max_05_topk_bitonic