Index
Run index
76 runs · 72 scored · 0 no perf · 4 fail · 0 err
methodology + notes
One row per (model, problem) cell, scored rows sorted by peak fraction desc. Click any row to open the full agent transcript on HuggingFace — every tool call, every reasoning step, the model's solution.py, and the result. Transcripts are published as per-run JSONL in the kernelbench-hard-traces dataset.