Skip to content
ISEGORIABenjamin Haire

Research note · Expanded edition · Visualizations updated 10 September 2026

Coding performance, cost and waiting

I put GPT-6 Astra, GPT-5.6 Sol and Claude Fable 5.1 side by side across selected reasoning levels. For this expanded edition I have brought together Astra medium/high/xhigh, Sol xhigh/max and Fable extra-high (xhigh), alongside the full Astra/Fable effort sweep.

Read the seven-page PDF (9 September edition)Six-way comparisonFull effort curvesSources and data

Astra medium is where I would start for interactive work. High improves the displayed terminal score, but its initial wait is much longer. Astra xhigh leads terminal coding, while Fable xhigh leads scientific coding in this six-way selection, with substantially higher mixed-suite token use and cost. I offer these as observations to trial against your own work, not as a universal best setting I claim to have found.

The six-way comparison

10 September 2026 snapshot · Fable uses adaptive reasoning with default fallback.

Astra high adds five terminal percentage points over medium for approximately 12% more mixed-suite cost and 17.2 times the first-answer wait. Fable xhigh exceeds Astra high by one rounded terminal point and six SciCode points, at 3.48 times the mixed-suite cost. I do not read a one-point lead as significant. Astra xhigh reaches 60% terminal and 56% SciCode, with 17k output tokens, $2.31 mixed-suite cost and a 208.73-second first-answer wait.

Read the six-way values as a table
10 September selected configurations
ConfigurationTerminalSciCodeOutputReasoning subsetUSD / taskFirst answer
Astra medium49%54%10k3k$1.545.42s
Astra high54%55%12k5k$1.7293.34s
Astra xhigh60%56%17k9k$2.31208.73s
Sol xhigh25%57%20k10k$1.1854.15s
Sol max40%57%29k17k$1.99140.25s
Fable 5.1 xhigh55%61%61k34k$5.98138.38s

The full Astra/Fable reasoning ladder

Original 8 September 2026 snapshot · Low, medium, high, xhigh and max.

Both terminal curves I observed peak at xhigh; Fable's scientific coding score continues to rise at max. Effort is an ordinal setting, so I drew the connecting lines as visual guides rather than a continuous fitted scaling law. I swept only Astra and Fable here: Sol appears at the two settings above and nowhere else.

How I read the evidence

Keep the workloads separate. Terminal-Bench v4.0 and SciCode are coding evaluations. Cost and token figures are weighted Intelligence Index v4.3 averages across a broader suite. First-answer latency comes from separate API performance prompts; it is not elapsed time to finish a tested patch. Reasoning tokens are already included in total output.

I date the snapshots deliberately. Fable xhigh's first-answer figure is 99.86 seconds in the 8 September sweep and 138.38 seconds in the 10 September comparison. I do not read that as a measured regression. Latency varies, and these samples cannot establish a stable completion-time multiplier.

The PDF I wrote covers per-token pricing for Astra/Fable, marginal effort trade-offs, methodology, limitations and a paired evaluation protocol I propose. I ran no new paid model trials. API costs do not translate directly into subscription usage allowances.

Sources and reusable data

For the selected comparison I used Artificial Analysis readings retrieved on 10 September. In the full sweep I have preserved the original 8 September source data and my reading of it at the time.