Small fixes and scoped subagent leaves.
Software engineering · September 2026
Choose the model for the shape of the work.
Claude Opus 5, Claude Fable 5, and every GPT-5.6 tier and effort level - compared without pretending vendor benchmarks form a complete league table.
Measured evidence
Published Coding Agent Index
High-end configurations reported in OpenAI's shared Artificial Analysis table. Opus 5 was not included.
Operating guide
Reasoning effort changes the job profile.
Routine features and test repair.
Ambiguous bugs and multi-file work.
Large migrations and deep verification.
Deepest single-agent reasoning.
Parallel research, implementation, and audit.
Sol and Terra support Ultra in the current Codex lineup. Luna, Opus 5, and Fable 5 stop at max.
Ultra under the microscope
More agents change the topology.
OpenAI reports Sol rising from 88.8% at max to 91.9% at Ultra on Terminal-Bench 2.1: a 3.1 percentage-point improvement.
Use Ultra when branches can proceed independently and one integrator owns interfaces, conflicts, and final proof. Stay on max for tightly coupled hotspots.
Routing guide
What to start with
Primary sources