Model · xAI

Grok 4.7

2 bench decks · 5/5 problems correct on canonical boards · 5 audited cells.

methodology + notes

How to read. Cell scores are peak fraction of the board roofline (Hard / most CUDA), geomean milliseconds for CUDA Native Sparse Attention, or speedup vs the torch baseline (Mega). Each cell has an unlimited agent budget; interrupted or overlapping resumes are documented in the audit. Audit chips come from the human/subagent reward-hack review of every published cell; scores from flagged sessions render dimmed — they don't count toward the charts.

Board summary bars are each score relative to the best published model on that board (1.00 = board leader); the printed number is the bench-native score.

Methodology. Rank per bench: valid passes (audited-clean correct cells / problems) desc, Cells with no audit annotation are shown but never scored. then mean normalized performance over the FULL problem deck (cell score / board best per problem; fail/invalid/missing cells count as 0) desc. Hack badge = flagged audited sessions / total audited sessions for that model; flagged = annotation verdict reward_hack | contamination | rubric_leak | suspect, or megakernel_authentic false (mega). Verdicts come from per-run audit YAMLs, not static lint. Hack rate is displayed, never a sort key. Browse the run index for transcripts, submitted solutions, checks, timing, and costs.

Board summarybars = share of each board's best model · numbers = bench-native score
Mega
6.38x1/1
CUDA
22.0%4/4
Megagrok
1/1 pass1 audited

RTX PRO 6000· canonical board

Kimi-Linear Decodepass
clean
6.38xfull-model speedup vs torch
1140 tok/s2048 ctx 6.33x8192 ctx 6.32x16384 ctx 6.50xptx

all audited grok attempts

Every attempted cell stays visible, including correctness failures, hardware mismatches, contaminated runs, and audit rejects. Only publishable results contribute to the board above.

RTX PRO 6000 Blackwell · Kimi-Linear Decodeclean
6.38xpublishable

Grok 4.7 built a genuine single-launch cooperative megakernel for the Kimi-Linear decode block in 63 minutes: in-kernel int4 dequant-GEMVs, short conv, gated-delta recurrence, absorbed MLA with a grid-parallel softmax, and the full 64-expert MoE, all behind one `cudaLaunchCooperativeKernel`. torch.profiler confirms 1.00 launches per step. Correctness is perfect (1.0000 on output, KDA state and MLA cache across all six check cases) and the templates are untouched. It is simply slow for this problem - 5.65x contended - and it stopped after one official benchmark.py run.

audited run
CUDAgrok · xhigh
4/4 pass4 audited

RTX PRO 6000· canonical board

GLM-5.2 Fused MoEpass
9.5%interesting
session 1h 41m
DeepSeek NSApass
0.984 msclean
session 59m
MegaQwen Decodepass
4.5%clean
session 1h 31m
Grid + MinGRU SPSpass
64.0%interesting
session 53m

all audited grok attempts

Every attempted cell stays visible, including correctness failures, hardware mismatches, contaminated runs, and audit rejects. Only publishable results contribute to the board above.

RTX_PRO_6000 · GLM-5.2 Fused MoEinteresting
9.48%publishable

Grok 4.7 wrote real fused-MoE plumbing in CUDA - histogram, expert-grouped gather with per-expert padding, SiLU-and-multiply, weighted scatter-reduce, and a hand-written decode GEMV - but every routed and shared GEMM is `cublasGemmStridedBatchedEx` (solution.py:297-301), not authored MMA. Same class as the published grok-4.6 (0.0939) and deepseek-v4-pro (0.0968) cells on this problem, both annotated `interesting` for the same reason: cuBLAS is not on the forbidden list, but the peak measures a library GEMM. Routing is taken as given, the shared expert always fires, every GEMM runs bf16-in / fp32-accum at full K and N on the real weights, nothing is precomputed outside the timed window, and the correctness margin is 13x inside the gate.

audited run
RTX_PRO_6000 · DeepSeek NSAclean
0.984 mspublishable

Grok 4.7 hand-wrote an SM120 NSA kernel in CUDA C++ with nvcuda::wmma bf16 m16n16k16 fragments, no cuBLAS and no torch op in the timed path: a block-sum kernel, a fused fp32 block-scoring plus top-8 kernel, a WMMA tile attention kernel for D=64 short sequences and a register-resident warp-per-query kernel for the long shapes and D=128. It computes the real op on every graded shape, verified against the reference at S=8192 and S=8191. It is also slow: 0.1002 of dense bf16 peak, against 0.5019 for the DeepSeek V4.1 Flash PTX kernel on the same problem.

audited run
RTX_PRO_6000 · MegaQwen Decodeclean
4.55%publishable

Grok 4.7 hand-wrote the whole Qwen3-0.6B decode block in CUDA C++ with inline PTX cache-hint loads: seven __global__ kernels (noise mix, fused QKV GEMV, split-K GQA scan, partial merge, O projection, fused SwiGLU gate/up, down projection), no cuBLAS, no torch op and no SDPA anywhere in the timed path. Every decode step reads the entire causal prefix at the reference's own bf16 KV round-trip precision, with fp32 accumulation. Nothing is windowed, strided, memoized or keyed to the graded context lengths. The design is launch-bound rather than bandwidth-bound at the short shapes: 25 kernel launches per token, which is why 0.0454 trails grok-4.6's 0.0542 and Opus 5's 0.0655 on the same problem.

audited run
RTX_PRO_6000 · Grid + MinGRU SPSinteresting
63.98%publishable

Grok 4.7 built a CUDA-graph rollout in 53 minutes and scored 0.6542: an encode kernel, a cuBLAS fp16 tensor-core GEMM for the three gate matrices, a fused MinGRU highway epilogue that keeps the hidden state in fp16, and greedy action plus LCG food respawn, the whole horizon replayed from one graph. The headline question is the fp16: problem.yaml declares `precision: fp32` and the entire graded policy runs in half. It is not a blind downgrade. Nine minutes in, before writing a kernel, the agent ran a controlled fp32-vs-bf16-vs-fp16 rollout at graded shapes, measured bf16 flipping positions and fp16 not, and picked fp16 on that evidence (transcript.jsonl:233). My independent strict oracle agrees at every shape and every seed benchmark.py grades: positions exact, rewards bitwise, logits 2.3e-6 against a 1e-3 gate. There is no shape branch on precision, no torch global is mutated, and check.py's own run(128,8) exercises the identical fp16 kernels. What the deck cannot see is the numeric-stress gate: solution.policy_forward is byte-identical PyTorch to reference.policy_forward, so the 1e-6 stress case measures exactly 0.0 and validates nothing about the kernels - the agent said so in its final message.

audited run