Model · DeepSeek
DeepSeek V4.1 Flash
3 bench decks · 11/11 problems correct on canonical boards · 11 audited cells.
methodology + notes
How to read. Cell scores are peak fraction of the board roofline (Hard / most CUDA), geomean milliseconds for CUDA Native Sparse Attention, or speedup vs the torch baseline (Mega). Each cell has an unlimited agent budget; interrupted or overlapping resumes are documented in the audit. Audit chips come from the human/subagent reward-hack review of every published cell; scores from flagged sessions render dimmed — they don't count toward the charts.
Board summary bars are each score relative to the best published model on that board (1.00 = board leader); the printed number is the bench-native score.
Methodology. Rank per bench: valid passes (audited-clean correct cells / problems) desc, Cells with no audit annotation are shown but never scored. then mean normalized performance over the FULL problem deck (cell score / board best per problem; fail/invalid/missing cells count as 0) desc. Hack badge = flagged audited sessions / total audited sessions for that model; flagged = annotation verdict reward_hack | contamination | rubric_leak | suspect, or megakernel_authentic false (mega). Verdicts come from per-run audit YAMLs, not static lint. Hack rate is displayed, never a sort key. Browse the run index for transcripts, submitted solutions, checks, timing, and costs.
RTX PRO 6000· canonical board
all audited deepseek-claude attempts
Every attempted cell stays visible, including correctness failures, hardware mismatches, contaminated runs, and audit rejects. Only publishable results contribute to the board above.
Hand-written Triton fp8 GEMM: tl.dot on two e4m3 operands lowering to mma.sync.m16n8k32 with an fp32 accumulator, per-output-channel dequant in the epilogue. 578.9 TFLOPS on 4096-cubed is above this part's 500 TFLOPS bf16 dense peak, which is independent physical proof the fp8 pipe is really running rather than a bf16-upcast fallback.
audited runTwo-kernel Triton chunk-parallel KDA forward. The intra kernel builds the decay-factored Gram matrix and inverts (I+L) exactly through a blocked product form on 16x16 diagonal blocks; the chain kernel walks the inter-chunk recurrence term-for-term against the reference. No CUTLASS needed: the prompt permits Triton and FLA is not even installed on the box.
audited runHand-written CUDA split-KV paged decode through load_inline: one launch, four warps merging (m, l, acc), an atomic ticket electing the last CTA to combine the chunk partials. What makes it interesting is the ceiling: the agent derived a 1205 GB/s bandwidth roof from torch.Tensor.sum() and stopped, 26% short of what this GPU delivers.
audited runSingle-launch CUDA bitonic selection, not a sort. Every element is packed into one 64-bit (monotonic key << 32) | index word, so values come back bit-exact and indices cannot duplicate. The merge compares a pair's second list backwards where it lies instead of copying it reversed into a gap, which makes the union bitonic for free.
audited runTriton variable-length grouped GEMM with SwiGLU genuinely fused into the epilogue. W_gate and W_up are interleaved column-wise into one (E, H, 2I) tensor so both projections come out of a single tl.dot into one accumulator, then reshape/split/silu is a register rename rather than a shared-memory round trip. 387.8 TFLOPS real, 77% of this part's bf16 dense peak.
audited runReal fused W4A16 Triton kernel. Packed uint8 is read inside the K loop and both nibble planes are widened, zero-shifted and scaled in registers, so no dequantised weight ever reaches memory and no pre-dequant cache exists. The scale is folded into the weight operand so one fp32 accumulator spans the whole K, measured at 2.9x the two-accumulator form.
audited runRTX PRO 6000· canonical board
all audited deepseek-claude attempts
Every attempted cell stays visible, including correctness failures, hardware mismatches, contaminated runs, and audit rejects. Only publishable results contribute to the board above.
DeepSeek V4.1 Flash built a real single-launch cooperative megakernel for the Kimi-Linear decode block: int4 dequant-GEMVs in registers, short conv, gated-delta recurrence, absorbed MLA, full MoE. Fast and honestly earned, but it computes the wrong answer inside the 0.98 cosine gate: a persistent atomicAdd accumulator is never zeroed, so every KDA layer inherits the previous one's activations.
audited runRTX PRO 6000· canonical board
all audited deepseek-claude attempts
Every attempted cell stays visible, including correctness failures, hardware mismatches, contaminated runs, and audit rejects. Only publishable results contribute to the board above.
DeepSeek V4.1 Flash hand-rolled the whole GLM-5.2 MoE layer in inline PTX: mma.sync bf16 tensor cores, ldmatrix, xor-swizzled cp.async, hist/scan/scatter token packing, one code path for every T. Every expert GEMM runs on real weights at full K and N; row padding makes it burn 25-50% more FLOPs than the roofline charges. Error margin 2.5x inside the gate.
audited runDeepSeek V4.1 Flash hand-wrote an SM120 NSA kernel in inline PTX: mma.sync bf16, ldmatrix, xor-swizzled cp.async, and an fp32 block-scoring top-8 prologue fused into the attention kernel. It reproduces the reference's exact tie-break and then over-computes, executing about 78% of the dense causal block triangle where the semantics need 14%.
audited runDeepSeek V4.1 Flash wrote a real SM120 cooperative megakernel for the 4-layer Qwen3-0.6B decode block: one cooperative launch per call with the step and layer loops inside, five grid.sync phases per layer, full-K warp GEMVs, and chunked GQA that reads every causal position on every step. Nothing windowed, sampled, or memoized.
audited runDeepSeek V4.1 Flash wrote a hand-rolled fp32 SM120 rollout: a cp.async double-buffered register-tile GEMM with the 3-layer MinGRU highway fused into the accumulators, the whole horizon replayed from a CUDA graph. It reproduces the reference environment down to the batch-global any-hit RNG gating that two other cells on this problem shortcut.
audited run