Physics Research Grade problems are research-level physics and mathematics derivations, each reconstructed from a recent arXiv preprint and scored only on whether a model reaches the exact terminal result — a direct read on reasoning ability rather than style.
Pass@1
Strict — whole problem solved
Performance
100%75%50%25%0%
32%
14%
14%
14%
9%
5%
5%
GPT-5.6-sol
Qwen 3.8 Max Preview
Claude Opus 4.8
Grok 4.5
Gemini 3.6 Flash
GLM 5.2
Nemotron 3 Ultra
Rubric score
Weighted partial credit
Each problem is broken into weighted subquestions, and the rubric score is the fraction of that credit an answer earns. It gives a model credit for correct structure and intermediate results even when it misses the final answer — where pass@1 above demands the complete result, this is the softer, partial view.
Performance
100%75%50%25%0%
51%
38%
30%
29%
26%
23%
20%
GPT-5.6-sol
Qwen 3.8 Max Preview
Claude Opus 4.8
Gemini 3.6 Flash
Grok 4.5
GLM 5.2
Nemotron 3 Ultra
Cost vs performance
Avg $ per task · pass@1
Each mark is one model, placed by the selected x-metric (log scale) against its strict pass@1 — cheaper or leaner to the left, stronger upward, so the top-left is the efficient frontier. Hover a mark for its full figures. Qwen's cost went untracked on its free route, so it is omitted from the cost view but still appears under the token axes.