Skip to content
ISEGORIABenjamin Haire

AI & software engineering · Interactive

Model Benchmark Atlas

Astra, Fable 5.1 and Opus 5: how reasoning effort changes measured quality, cost and response time, and which tests matter for game development.

01 / Compare

What more reasoning buys, and what it costs

Pick a metric and an effort setting. The main chart follows one metric; the panels below show all fifteen at once, and the last chart sets quality against cost or waiting time.

Reasoning effort

    Equal effort names do not mean equal compute. Values are dated measurements, not live results.

    Vertical scale adapts to each metric. Lines connect five discrete settings; intermediate effort levels were not measured. Hover or use the arrow keys to read a setting; click to select it.

    Every metric at once

    Each panel has its own vertical scale. Hover a panel to read the same effort across all fifteen; select one to open it in the main chart.

    Quality

    Cost and time

    Quality against cost

    Each model traces its five effort settings from low to max. Up and to the left is better. The dashed step line joins the frontier: configurations that no other plotted configuration beats on both axes at once. Values are rounded and from one snapshot, so near-ties are not meaningful.

    Against
    • Frontier
    • Configuration on the frontier
    • Larger mark: selected effort

    Exact values
    MetricAstraFable 5.1Opus 5

    Token prices and actual task cost

    Standard API rates in USD per million tokens, recorded on 15 September 2026. Subscription plans are separate. Token prices are inputs to a bill; the cost-per-task chart also reflects how many tokens each evaluation consumes.

    ModelInputCache readOutput
    Astra$10$1.00$50
    Fable 5.1$10$0.25$50
    Opus 5$5$0.50$25

    Astra rates shown for inputs up to 272k tokens. Cache writes, tool fees, long-context tiers and other service modes may differ. Consult the linked provider documentation before budgeting.

    Astra ↗ · Fable ↗ · Opus ↗

    02 / What the tests measure

    What is actually being tested?

    A benchmark is a task definition plus a test environment and a scoring rule. Similar-looking numbers may measure very different abilities.

    SciCode: scientific knowledge expressed as working code

    The full benchmark contains 80 main scientific problems divided into 338 subproblems. Artificial Analysis uses the 288 test-set subproblems, supplies scientist-written background, and scores Python code through unit tests. The measured skill combines understanding a scientific specification, translating it into an algorithm, and implementing that algorithm correctly.

    Illustrative task, not a quoted benchmark item: implement a numerical routine with the right units, array shapes and boundary conditions. Explaining the physics beautifully would not pass if the function returns incorrect values.

    A 56% subproblem score does not mean the model completes 56% of research projects. It also does not directly measure C++, engine architecture, shader quality or frame time. The later SciCode-Verified audit reports defects in the original benchmark; the chart here retains the published Artificial Analysis SciCode values rather than substituting results from the revised test.

    SciCode ↗ · SciCode-Verified ↗

    GDPval: professional work products, not a knowledge quiz

    GDPval-AA v2 evaluates 220 public tasks spanning 44 occupations. Agents use tools to produce files such as reports, spreadsheets and presentations. A panel of model judges compares submitted work; the reported Elo scale is anchored to human expert work at 1000. The v2 environment permits up to 250 turns.

    Illustrative task: read reference files, calculate a justified recommendation, and deliver a usable spreadsheet and memo. Factual knowledge helps, but file handling, following requirements, checking calculations and communicating the result also affect the outcome.

    Elo describes relative performance in this evaluation. A score of 1700 is not 70% better than a human and is not a percentage of jobs an agent can replace. The environment, tools and judging procedure are part of the result.

    GDPval-AA v2 ↗ · Methodology ↗

    Why can a flagship model lag on knowledge work?

    “Flagship” names a product position; it does not mean winning every test. At xhigh in these snapshots, Astra and Fable 5.1 both score 53 on the rounded overall index, while Opus 5 scores 50. On GDPval-AA, however, Opus 5 is ahead of Astra. Astra also leads Opus on GDP.pdf and AutomationBench-AA at that setting. Professional deliverables, document reasoning and workflow automation are distinct capabilities.

    These observations establish the pattern, not its internal cause. Training emphasis, tool behavior, effort allocation and the judging rubric are possible contributors; the published scores do not isolate which explains the gap. Choose by the work you need done, then verify on representative tasks.

    Terminal-Bench 4.0 · Can an agent finish a terminal task?

    Tool-driven tasks are checked by executable tests. The score describes successful task completion under the evaluation setup, rather than coding style or aesthetic judgment.

    Methodology ↗
    AA-Omniscience · Will it answer facts reliably?

    Open factual questions, with an index that accounts for wrong assertions. An index of 43 is not 43% accuracy. This is closer to factual reliability than the quality of a slide deck.

    Methodology ↗
    Humanity’s Last Exam · Can it solve very difficult academic questions?

    A multidisciplinary academic challenge. It is useful for hard reasoning, but answering an exam question and delivering a working software change require different behaviors.

    Methodology ↗
    GDP.pdf · Can it reason from a long professional document?

    The model answers questions grounded in a PDF. The displayed all-pass score requires the response to satisfy every assessed criterion. It is a different task and score from GDPval-AA.

    Methodology ↗
    AA-LCR v1.1 · Can it find and combine distant information?

    Long-context reasoning checks whether the model can use information spread across a large input. A large advertised context window does not guarantee good use of that context.

    Methodology ↗
    AutomationBench-AA · Can it carry out a business workflow?

    Tasks require actions through software APIs. This tests tool use and correct state changes; it is not a visual-design benchmark.

    Methodology ↗
    CritPt · Can it solve research-level physics problems?

    Physics problems stress scientific reasoning. This does not directly test whether an implementation is numerically stable, fast or well integrated into a game engine.

    Methodology ↗
    AA-Briefcase · Can it produce a useful business deliverable?

    Complex business work is assessed across task success, analysis and presentation. Its combined Elo rating cannot be compared numerically with GDPval-AA Elo as though they shared one scale.

    Methodology ↗

    03 / Game development

    Game development

    For engine and rendering work, start with tests that require a running project or measured GPU execution. A plausible explanation or attractive screenshot is only one part of the evidence.

    Engine behavior

    GameDevBench ↗

    333 Godot tasks

    Changes to scenes, shaders, sprites, animation, UI and gameplay. Relevant to visual iteration inside an engine; not a full native-renderer performance suite.

    GameEngineBench ↗

    110 Unreal C++ tasks

    Integrated C++ behavior, including networking and object lifecycle. Uses runtime tests and model-judge review. Performance and platform coverage are deliberately lighter.

    Visual construction

    BlenderGym ↗

    Code-based graphics editing

    Reconstruct and edit 3D graphics through code. Useful for scene and graphics manipulation; it does not establish real-time renderer engineering.

    Shader Frontier v1 ↗

    19 attempts · 2 targets

    Recorded shader-building outputs, elapsed time and cost. Each entry is a single attempt; repeat-run reliability and broad generalization are unestablished.

    GPU compute

    KernelBench ↗

    Correctness + speedup

    Generate GPU kernels for PyTorch workloads. The fast₁ measure counts correct kernels faster than the baseline. It tests AI compute, not full rasterization or game frame rate.

    SOL-ExecBench ↗

    235 CUDA optimization tasks

    Measures progress toward analytical hardware efficiency limits on NVIDIA Blackwell AI workloads. This is not direct evidence for Metal, a render graph or temporal image stability.

    The relevant published result we can show

    GameDevBench · percentage of 333 tasks passed · 0–100% scale

    Astra · high · Codex68.8%
    Fable 5 · xhigh · Claude Code67.3%

    95% confidence intervals: ±5.0 percentage points for both. Best multimodal feedback and agent setup per model; these are not identical configurations.

    Fable 5 is not Fable 5.1. Fable 5.1 and Opus 5 have no entries on this leaderboard as checked on 17 September 2026. The 1.5-point gap shown here is not a reliable basis for claiming an overall winner.

    GameDevBench ↗

    A practical renderer evaluation

    Suggested evaluation design, not an existing published score: give each model the same repository revision, task, hardware, tools and budget; repeat runs and retain failures.

    1. Correctness

      Build and runtime tests, reference images, camera movement, temporal artifacts and regression checks.

    2. Performance

      GPU milliseconds, median and slow-tail frame times, memory consumption and repeatable captures on target hardware.

    3. Delivery cost

      Wall-clock completion time, API spend and human repair effort, alongside whether the final change is usable.

    04 / Sources & limits

    How to read this atlas

    English and Japanese editions share the same data and controls.

    1. Artificial Analysis · methodology ↗
    2. SciCode · evaluation ↗
    3. GDPval-AA v2 · evaluation ↗
    4. SciCode-Verified · benchmark audit ↗
    5. GameDevBench · project and leaderboard ↗
    6. GameEngineBench · paper ↗
    7. BlenderGym · paper ↗
    8. Shader Frontier v1 · run collection ↗
    9. KernelBench · evaluation and code ↗
    10. SOL-ExecBench · paper ↗
    11. Astra · API model documentation ↗
    12. Fable · provider page ↗
    13. Opus 5 · provider announcement ↗
    14. Astra / Fable 5.1 · low ↗
    15. Astra / Fable 5.1 · medium ↗
    16. Astra / Fable 5.1 · high ↗
    17. Astra / Fable 5.1 · xhigh ↗
    18. Astra / Fable 5.1 · max ↗
    19. Opus 5 · low / max ↗
    20. Astra / Opus 5 · medium ↗
    21. Astra / Opus 5 · high ↗
    22. Astra / Opus 5 · xhigh ↗