writeups
long-form posts on design choices, rubric leaks, and methodology.
[ latest ]
Hard: Eight Scored Problems, Frontier Models, Two Rubric Leaks
One Blackwell GPU, a focused CUDA problem deck, real coding-agent CLIs as the harness. Frontier models swept; only GPT-5.5 xhigh solved every problem. Two problems leak the rubric: five models all took the same bf16 shortcut on FP8 GEMM, and the only model that implemented Kahan compensated summation scored lowest of the passes.