Our LLM is uniquely good at high-throughput data processing: video, audio, and live text streams, all understood in real time. Whether gaming, using software, or conducting research, our model always runs at competitive speed.
Video games are the ultimate stress test for real time video processing. Here Star reads the screen and reacts in real time, no game-specific scripting. See more at Star plays games and Star for robotics.
See how Star uses real software GUIs to get the job done.
Relevant benchmarks to show comparison:
Star can parse a live feed at inhuman speeds, noticing change, reacting fast enough that the answer still applies by the time the joint moves. See more at Star for robotics.
Sourcing, competitive analysis, and long web workflows done at inhuman speed, for pennies.
The full comparison, relative to published or in-house data. Full set at the benchmarks page.
| Eval | FrontierRL Star | Claude Fable 5 | Claude Fable 5 (Xhigh) | Claude Opus 4.8 | Claude Opus 5 | Claude Sonnet 5 | DeepSeek v4 Flash | DeepSeek v4 Pro | Gemini 3 Pro | Gemini 3.1 Pro | Gemini 3.1 Pro Preview | Gemini 3.6 Flash | GPT‑5 | GPT‑5.5 | GPT‑5.6 Luna | GPT‑5.6 Sol | GPT‑5.6 Sol (High) | GPT‑5.6 Sol (Max) | GPT‑5.6 Terra | Kimi K3 | Qwen3.7 Max | Source for other model's metrics |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BenchCAD (python tool) | 81.75% | — | — | 51.8% | — | — | — | — | — | — | — | — | — | 55.8% | 73.9% | 83.4% | — | — | 78.2% | — | — | openai.com/index/gpt-5-6 |
| IFEval | 97.97% | — | — | — | — | — | 91.6% | 93.4% | — | — | — | — | — | — | — | — | — | 96.64% | — | — | 94.2% | In-house tested by FrontierRL |
| OSWorld Verified | 85.36% | 85% | — | 83.4% | — | — | — | — | — | 76.2% | — | — | — | 78.7% | — | — | — | — | — | — | — | anthropic.com/claude/mythos |
| Terminal-Bench 2.1 | 75.28% | 80.52% | — | — | 84.64% | 74.53% | — | — | — | — | 70.79% | 73.78% | — | 76.40% | — | 85.77% | — | — | — | 80.90% | — | openai.com/index/gpt-5-6 |
| Video-MME v2 | 83% | — | 77% | — | — | — | — | — | 66.1% | — | — | — | 44.7% | — | — | — | 67% | 69% | — | — | — | In-house tested by FrontierRL |