GameDevBench ↗
Changes to scenes, shaders, sprites, animation, UI and gameplay. Relevant to visual iteration inside an engine; not a full native-renderer performance suite.
AI & software engineering · Interactive
Astra, Fable 5.1 and Opus 5: how reasoning effort changes measured quality, cost and response time, and which tests matter for game development.
01 / Compare
Pick a metric and an effort setting. The main chart follows one metric; the panels below show all fifteen at once, and the last chart sets quality against cost or waiting time.
Equal effort names do not mean equal compute. Values are dated measurements, not live results.
Each panel has its own vertical scale. Hover a panel to read the same effort across all fifteen; select one to open it in the main chart.
Each model traces its five effort settings from low to max. Up and to the left is better. The dashed step line joins the frontier: configurations that no other plotted configuration beats on both axes at once. Values are rounded and from one snapshot, so near-ties are not meaningful.
| Metric | Astra | Fable 5.1 | Opus 5 |
|---|
Standard API rates in USD per million tokens, recorded on 15 September 2026. Subscription plans are separate. Token prices are inputs to a bill; the cost-per-task chart also reflects how many tokens each evaluation consumes.
| Model | Input | Cache read | Output |
|---|---|---|---|
| Astra | $10 | $1.00 | $50 |
| Fable 5.1 | $10 | $0.25 | $50 |
| Opus 5 | $5 | $0.50 | $25 |
Astra rates shown for inputs up to 272k tokens. Cache writes, tool fees, long-context tiers and other service modes may differ. Consult the linked provider documentation before budgeting.
Astra ↗ · Fable ↗ · Opus ↗02 / What the tests measure
A benchmark is a task definition plus a test environment and a scoring rule. Similar-looking numbers may measure very different abilities.
The full benchmark contains 80 main scientific problems divided into 338 subproblems. Artificial Analysis uses the 288 test-set subproblems, supplies scientist-written background, and scores Python code through unit tests. The measured skill combines understanding a scientific specification, translating it into an algorithm, and implementing that algorithm correctly.
Illustrative task, not a quoted benchmark item: implement a numerical routine with the right units, array shapes and boundary conditions. Explaining the physics beautifully would not pass if the function returns incorrect values.
A 56% subproblem score does not mean the model completes 56% of research projects. It also does not directly measure C++, engine architecture, shader quality or frame time. The later SciCode-Verified audit reports defects in the original benchmark; the chart here retains the published Artificial Analysis SciCode values rather than substituting results from the revised test.
SciCode ↗ · SciCode-Verified ↗GDPval-AA v2 evaluates 220 public tasks spanning 44 occupations. Agents use tools to produce files such as reports, spreadsheets and presentations. A panel of model judges compares submitted work; the reported Elo scale is anchored to human expert work at 1000. The v2 environment permits up to 250 turns.
Illustrative task: read reference files, calculate a justified recommendation, and deliver a usable spreadsheet and memo. Factual knowledge helps, but file handling, following requirements, checking calculations and communicating the result also affect the outcome.
Elo describes relative performance in this evaluation. A score of 1700 is not 70% better than a human and is not a percentage of jobs an agent can replace. The environment, tools and judging procedure are part of the result.
GDPval-AA v2 ↗ · Methodology ↗“Flagship” names a product position; it does not mean winning every test. At xhigh in these snapshots, Astra and Fable 5.1 both score 53 on the rounded overall index, while Opus 5 scores 50. On GDPval-AA, however, Opus 5 is ahead of Astra. Astra also leads Opus on GDP.pdf and AutomationBench-AA at that setting. Professional deliverables, document reasoning and workflow automation are distinct capabilities.
These observations establish the pattern, not its internal cause. Training emphasis, tool behavior, effort allocation and the judging rubric are possible contributors; the published scores do not isolate which explains the gap. Choose by the work you need done, then verify on representative tasks.
Tool-driven tasks are checked by executable tests. The score describes successful task completion under the evaluation setup, rather than coding style or aesthetic judgment.
Methodology ↗Open factual questions, with an index that accounts for wrong assertions. An index of 43 is not 43% accuracy. This is closer to factual reliability than the quality of a slide deck.
Methodology ↗A multidisciplinary academic challenge. It is useful for hard reasoning, but answering an exam question and delivering a working software change require different behaviors.
Methodology ↗The model answers questions grounded in a PDF. The displayed all-pass score requires the response to satisfy every assessed criterion. It is a different task and score from GDPval-AA.
Methodology ↗Long-context reasoning checks whether the model can use information spread across a large input. A large advertised context window does not guarantee good use of that context.
Methodology ↗Tasks require actions through software APIs. This tests tool use and correct state changes; it is not a visual-design benchmark.
Methodology ↗Physics problems stress scientific reasoning. This does not directly test whether an implementation is numerically stable, fast or well integrated into a game engine.
Methodology ↗Complex business work is assessed across task success, analysis and presentation. Its combined Elo rating cannot be compared numerically with GDPval-AA Elo as though they shared one scale.
Methodology ↗03 / Game development
For engine and rendering work, start with tests that require a running project or measured GPU execution. A plausible explanation or attractive screenshot is only one part of the evidence.
Changes to scenes, shaders, sprites, animation, UI and gameplay. Relevant to visual iteration inside an engine; not a full native-renderer performance suite.
Integrated C++ behavior, including networking and object lifecycle. Uses runtime tests and model-judge review. Performance and platform coverage are deliberately lighter.
Reconstruct and edit 3D graphics through code. Useful for scene and graphics manipulation; it does not establish real-time renderer engineering.
Recorded shader-building outputs, elapsed time and cost. Each entry is a single attempt; repeat-run reliability and broad generalization are unestablished.
Generate GPU kernels for PyTorch workloads. The fast₁ measure counts correct kernels faster than the baseline. It tests AI compute, not full rasterization or game frame rate.
Measures progress toward analytical hardware efficiency limits on NVIDIA Blackwell AI workloads. This is not direct evidence for Metal, a render graph or temporal image stability.
GameDevBench · percentage of 333 tasks passed · 0–100% scale
95% confidence intervals: ±5.0 percentage points for both. Best multimodal feedback and agent setup per model; these are not identical configurations.
Fable 5 is not Fable 5.1. Fable 5.1 and Opus 5 have no entries on this leaderboard as checked on 17 September 2026. The 1.5-point gap shown here is not a reliable basis for claiming an overall winner.
GameDevBench ↗Suggested evaluation design, not an existing published score: give each model the same repository revision, task, hardware, tools and budget; repeat runs and retain failures.
Build and runtime tests, reference images, camera movement, temporal artifacts and regression checks.
GPU milliseconds, median and slow-tail frame times, memory consumption and repeatable captures on target hardware.
Wall-clock completion time, API spend and human repair effort, alongside whether the final change is usable.
04 / Sources & limits
English and Japanese editions share the same data and controls.