Where STAR stands today, relative to published or in-house data.
FrontierRL is constantly running more benchmarks — results land on this page as they finish.
Data source for other models: OpenAI, “GPT‑5.6” announcement — openai.com/index/gpt-5-6 (July 2026).
Figures recharted by FrontierRL; values reproduced as published.
| Eval | FrontierRL STAR | GPT‑5.6 Sol | GPT‑5.6 Sol Ultra | GPT‑5.6 Terra | GPT‑5.6 Luna | GPT‑5.5 | Claude Mythos 5 | Claude Mythos Preview | Claude Opus 4.8 | Gemini 3.1 Pro Preview | DeepSeek v4 Flash | DeepSeek v4 Pro | Qwen3.7 Max | GPT‑5.5 Medium |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BenchCAD (python tool) | 81.75% | 83.4% | — | 78.2% | 73.9% | 55.8% | 65% | 61% | 51.8% | — | — | — | — | — |
Prompt-level strict accuracy: IFEval (541 verifiable-instruction prompts), measured in-house by FrontierRL | ||||||||||||||
| IFEval | 97.97% | — | — | — | — | — | — | — | — | — | 91.6% | 93.4% | 94.2% | 96.2% |
FrontierRL STAR results for these land here as they finish.
| Eval | FrontierRL STAR | GPT‑5.6 Sol | GPT‑5.6 Sol Ultra | GPT‑5.6 Terra | GPT‑5.6 Luna | GPT‑5.5 | Claude Mythos 5 | Claude Mythos Preview | Claude Fable 5 | Claude Opus 4.8 | Gemini 3.1 Pro Preview |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Terminal-Bench 2.1 | — | 88.8% | 91.9% | 87.4% | 84.7% | 85.6% | 88% | — | 83.1% | 78.9% | 70.7% |
| MMMU Pro (with tools) | — | 84.6% | — | 82% | 79.5% | 83.2% | — | — | — | — | — |
| GPQA Diamond | — | 94.6% | — | 92.9% | 92.3% | 93.6% | 94.1% | 94.6% | 92.6% | 92% | 94.3% |
| Toolathlon | — | 58% | — | 53.1% | 53.4% | 55.6% | 61.7% | 61.1% | 61.7% | 59.9% | 48.8% |
| GraphWalks BFS 256k f1 | — | 90.7% | — | 76.9% | 81.3% | 73.7% | 91.1% | 85.7% | — | 85.9% | — |
| HealthBench Professional | — | 60.5% | — | 57.7% | 55.7% | 49.5% | — | — | 60.9% | 53% | — |
| DeepSWE v1.1 | — | 72.7% | — | 69.6% | 67.2% | 67% | — | — | 69.7% | 59% | 11.8% |