Aeon AI Risk Management
Model Economics Lab
Start with all five new releases plus measured shortlist anchors for GPT-5.6 Sol, Claude Opus 5, Kimi K3, and GLM-5.2, then expand into popular, value, and historical views across all 40 retained operating points.
Executive finding
Buy reasoning effort by task class, not by brand. Opus 5 xhigh leads the measured field, Sol max is nearly tied at lower cost, Grok high is a strong upper-tier value point, and Muse Spark strengthens the middle of the curve.
Latest release evidence
Grok 4.6, Gemini 3.7 Flash, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, and Muse Spark 1.2 all appear as provisional diamonds in the default plot. Qwen3.8 remains evidence-only until defensible score, cost, and token coordinates are public.
CursorBench 3.2 cross-check
CursorBench is presented as a separate benchmark lens with its own score, cost, total-token, and agent-step measurements. Grok 4.6 and Gemini 3.7 Flash also appear as cross-suite diamonds in the opening plot, while Grok 4.5 and GPT-5.5 remain historical anchors.
Master plot
The horizontal axis is published or explicitly proxied task cost, the vertical axis is each point's published coding score in its native suite, and marker area represents published or proxied tokens. Filled circles are comparable measurements, diamonds are provisional cross-suite or proxy points, and hollow circles are modeled.
Reader views
Current releases, Popular models, Value challengers, and All / history separate the nine-point default from the complete retained dataset. Users can then expand all effort levels, filter evidence, toggle trajectories, and change the cost scale.
Measured anchors
Opus 5 xhigh reaches 66.74 at $8.23 per task. Sol max reaches 66.57 at $7.08. Grok 4.5 high reaches 64.44 at $2.59. Kimi K3 reaches 61.34 at $3.18. Muse Spark 1.1 reaches 53.54 at $1.43. Luna high reaches 51.42 below $1.
Effort economics
Higher effort creates nonlinear token and cost growth. The last capability points can be valuable for ambiguous, long-horizon, or high-consequence work, but are usually inefficient for routine coding.
Modeled points
Where no comparable public run exists, Aeon uses neighboring model-family effort curves or a vendor-published effort curve. Every estimate is marked as modeled and is not presented as a benchmark result.
Limitations
Output tokens do not include every input, cache, retry, tool, and harness cost. Coding-agent token use can vary substantially by task and harness, so organizations should validate routing against their own repositories and acceptance tests.
Downloadable dataset
The full operating-point dataset is available in CSV and JSON with model, provider, effort, published score, native score metric, task cost and basis, tokens and token metric, evidence status, modeling basis, and source URL.
Dataset license and attribution
Aeon publishes this research compilation for evaluation and citation. Cite Aeon and link to this page when referencing the compilation. Underlying benchmark data and source materials remain subject to their publishers' terms. This notice does not relicense third-party material. Contact info@airiskmanagement.ca for redistribution or commercial reuse.
Update policy
Aeon adds every new release to the default plot, uses diamonds for incomplete or mixed-basis coordinates, reviews the measured shortlist monthly, and retains older observations for audit and reader demand.
Operating recommendation
Use a routing ladder that assigns cheaper effort to bounded tasks and reserves frontier settings for ambiguity, long dependency chains, and changes with high failure cost.