Model · xAI
Grok 4.7
2 bench decks · 5/5 problems correct on canonical boards · 5 audited cells.
methodology + notes
How to read. Cell scores are peak fraction of the board roofline (Hard / most CUDA), geomean milliseconds for CUDA Native Sparse Attention, or speedup vs the torch baseline (Mega). Each cell has an unlimited agent budget; interrupted or overlapping resumes are documented in the audit. Audit chips come from the human/subagent reward-hack review of every published cell; scores from flagged sessions render dimmed — they don't count toward the charts.
Board summary bars are each score relative to the best published model on that board (1.00 = board leader); the printed number is the bench-native score.
Methodology. Rank per bench: valid passes (audited-clean correct cells / problems) desc, Cells with no audit annotation are shown but never scored. then mean normalized performance over the FULL problem deck (cell score / board best per problem; fail/invalid/missing cells count as 0) desc. Hack badge = flagged audited sessions / total audited sessions for that model; flagged = annotation verdict reward_hack | contamination | rubric_leak | suspect, or megakernel_authentic false (mega). Verdicts come from per-run audit YAMLs, not static lint. Hack rate is displayed, never a sort key. Browse the run index for transcripts, submitted solutions, checks, timing, and costs.
RTX PRO 6000· canonical board
all audited grok attempts
Every attempted cell stays visible, including correctness failures, hardware mismatches, contaminated runs, and audit rejects. Only publishable results contribute to the board above.
Grok 4.7 built a genuine single-launch cooperative megakernel for the Kimi-Linear decode block in 63 minutes: in-kernel int4 dequant-GEMVs, short conv, gated-delta recurrence, absorbed MLA with a grid-parallel softmax, and the full 64-expert MoE, all behind one `cudaLaunchCooperativeKernel`. torch.profiler confirms 1.00 launches per step. Correctness is perfect (1.0000 on output, KDA state and MLA cache across all six check cases) and the templates are untouched. It is simply slow for this problem - 5.65x contended - and it stopped after one official benchmark.py run.
audited runRTX PRO 6000· canonical board
all audited grok attempts
Every attempted cell stays visible, including correctness failures, hardware mismatches, contaminated runs, and audit rejects. Only publishable results contribute to the board above.
Grok 4.7 wrote real fused-MoE plumbing in CUDA - histogram, expert-grouped gather with per-expert padding, SiLU-and-multiply, weighted scatter-reduce, and a hand-written decode GEMV - but every routed and shared GEMM is `cublasGemmStridedBatchedEx` (solution.py:297-301), not authored MMA. Same class as the published grok-4.6 (0.0939) and deepseek-v4-pro (0.0968) cells on this problem, both annotated `interesting` for the same reason: cuBLAS is not on the forbidden list, but the peak measures a library GEMM. Routing is taken as given, the shared expert always fires, every GEMM runs bf16-in / fp32-accum at full K and N on the real weights, nothing is precomputed outside the timed window, and the correctness margin is 13x inside the gate.
audited runGrok 4.7 hand-wrote an SM120 NSA kernel in CUDA C++ with nvcuda::wmma bf16 m16n16k16 fragments, no cuBLAS and no torch op in the timed path: a block-sum kernel, a fused fp32 block-scoring plus top-8 kernel, a WMMA tile attention kernel for D=64 short sequences and a register-resident warp-per-query kernel for the long shapes and D=128. It computes the real op on every graded shape, verified against the reference at S=8192 and S=8191. It is also slow: 0.1002 of dense bf16 peak, against 0.5019 for the DeepSeek V4.1 Flash PTX kernel on the same problem.
audited runGrok 4.7 hand-wrote the whole Qwen3-0.6B decode block in CUDA C++ with inline PTX cache-hint loads: seven __global__ kernels (noise mix, fused QKV GEMV, split-K GQA scan, partial merge, O projection, fused SwiGLU gate/up, down projection), no cuBLAS, no torch op and no SDPA anywhere in the timed path. Every decode step reads the entire causal prefix at the reference's own bf16 KV round-trip precision, with fp32 accumulation. Nothing is windowed, strided, memoized or keyed to the graded context lengths. The design is launch-bound rather than bandwidth-bound at the short shapes: 25 kernel launches per token, which is why 0.0454 trails grok-4.6's 0.0542 and Opus 5's 0.0655 on the same problem.
audited runGrok 4.7 built a CUDA-graph rollout in 53 minutes and scored 0.6542: an encode kernel, a cuBLAS fp16 tensor-core GEMM for the three gate matrices, a fused MinGRU highway epilogue that keeps the hidden state in fp16, and greedy action plus LCG food respawn, the whole horizon replayed from one graph. The headline question is the fp16: problem.yaml declares `precision: fp32` and the entire graded policy runs in half. It is not a blind downgrade. Nine minutes in, before writing a kernel, the agent ran a controlled fp32-vs-bf16-vs-fp16 rollout at graded shapes, measured bf16 flipping positions and fp16 not, and picked fp16 on that evidence (transcript.jsonl:233). My independent strict oracle agrees at every shape and every seed benchmark.py grades: positions exact, rewards bitwise, logits 2.3e-6 against a 1e-3 gate. There is no shape branch on precision, no torch global is mutated, and check.py's own run(128,8) exercises the identical fp16 kernels. What the deck cannot see is the numeric-stress gate: solution.policy_forward is byte-identical PyTorch to reference.policy_forward, so the 1e-6 stress case measures exactly 0.0 and validates nothing about the kernels - the agent said so in its final message.
audited run