FrontierRL
Book an intro

Benchmarks.

Where STAR stands today, relative to published or in-house data.
FrontierRL is constantly running more benchmarks — results land on this page as they finish.

EvalFrontierRL StarClaude Fable 5Claude Fable 5 (Xhigh)Claude Opus 4.8Claude Opus 5Claude Sonnet 5DeepSeek v4 FlashDeepSeek v4 ProGemini 3 ProGemini 3.1 ProGemini 3.1 Pro PreviewGemini 3.6 FlashGPT‑5GPT‑5.5GPT‑5.6 LunaGPT‑5.6 SolGPT‑5.6 Sol (High)GPT‑5.6 Sol (Max)GPT‑5.6 TerraKimi K3Qwen3.7 MaxSource for other model's metrics
BenchCAD (python tool)81.75%51.8%55.8%73.9%83.4%78.2%openai.com/index/gpt-5-6
IFEval97.97%91.6%93.4%96.64%94.2%In-house tested by FrontierRL
OSWorld Verified85.36%85%83.4%76.2%78.7%anthropic.com/claude/mythos
Terminal-Bench 2.175.28%80.52%84.64%74.53%70.79%73.78%76.40%85.77%80.90%openai.com/index/gpt-5-6
Video-MME v283%77%66.1%44.7%67%69%In-house tested by FrontierRL

Dropping Soon!

FrontierRL Star results for these land here as they finish.

EvalFrontierRL StarClaude Fable 5Claude Opus 4.8GPT‑5.5GPT‑5.6 LunaGPT‑5.6 SolGPT‑5.6 Terra
DeepSWE v1.169.7%59%67%67.2%72.7%69.6%
GPQA Diamond92.6%92%93.6%92.3%94.6%92.9%
GraphWalks BFS 256k f185.9%73.7%81.3%90.7%76.9%
MMMU Pro (with tools)83.2%79.5%84.6%82%
Toolathlon61.7%59.9%55.6%53.4%58%53.1%

See Star in action