Three releases frame this week's coding-model economics.
Claude Opus 5 now holds the highest measured point in the Aeon Model Economics Lab. Meta Muse Spark 1.1 does not match frontier capability, but it reaches a useful coding score at a much lower task cost. Alibaba's Qwen3.8-Max-Preview is available to test, but it does not yet have the comparable benchmark and pricing evidence required for a defensible point on the chart.
That distinction matters. Model selection is not a release-date contest. It is a routing decision across capability, completed-task cost, token footprint, and evidence quality.
What changed in the master plot
We refreshed the Aeon Model Economics Lab against the current Artificial Analysis Coding Agent Index v1.3 data.
The plot keeps three variables visible:
- Vertical axis: coding capability on the Artificial Analysis Coding Agent Index.
- Horizontal axis: completed-task cost in United States dollars.
- Bubble area: output tokens emitted per task.
Every plotted point also carries the model and reasoning-effort label. Filled circles are measured in the same benchmark framework. Hollow circles are explicit estimates where a public effort-level measurement is missing.
| Model and setting | Coding Agent Index | Cost per task | Output tokens | Evidence |
|---|---|---|---|---|
| Claude Opus 5 xhigh | 66.74 | USD 8.23 | 72.8k | Measured |
| GPT-5.6 Sol max | 66.57 | USD 7.08 | 54.9k | Measured |
| Claude Fable 5 max | 65.85 | USD 11.71 | 73.6k | Measured, fallback configuration |
| Grok 4.5 high | 64.44 | USD 2.59 | 40.0k | Measured |
| Kimi K3 | 61.34 | USD 3.18 | 88.8k | Measured |
| Muse Spark 1.1 xhigh | 53.54 | USD 1.43 | 34.0k | Measured |
| GPT-5.6 Luna high | 51.42 | USD 0.96 | 32.3k | Measured |
| Qwen3.8-Max-Preview | Pending | Pending | Not measured | Preview watch |
Historical scores on the page changed because the benchmark is now on version 1.3. That is a harness refresh, not evidence that older models suddenly lost intelligence.
Opus 5 leads, but max is not the best Opus setting
Claude Opus 5 xhigh reaches 66.74 at USD 8.23 per task with 72.8k output tokens. That is the highest measured Coding Agent Index point in the refreshed dataset.
The more important operating insight is inside the Opus effort curve. Opus 5 max costs USD 8.95, emits 80.5k output tokens, and scores 65.53. In this run, xhigh is both cheaper and more capable.
That is a useful warning against defaulting every difficult task to the highest available reasoning setting. More test-time effort can increase cost and latency without improving the completed result. Production routing should use the lowest setting that reliably clears the task's acceptance bar.
Artificial Analysis's Opus 5 analysis also reports five effort levels and provider pricing of USD 5 per million input tokens and USD 25 per million output tokens. The completed-task points are more useful than token list prices because they include the behavior of the model and harness on actual coding work.
Sol is almost tied at a lower task cost
GPT-5.6 Sol max reaches 66.57 at USD 7.08 per task with 54.9k output tokens. It trails Opus 5 xhigh by only 0.17 index points while costing about USD 1.15 less per benchmark task and emitting fewer output tokens.
That does not prove Sol is better for every repository. It does mean a buyer should not pay for Opus 5 by default based on the top score alone. The two points are close enough that repository-specific acceptance tests, tool reliability, latency, and review burden should decide the route.
Grok remains the upper-tier value point
Grok 4.5 high reaches 64.44 at USD 2.59 per task. It gives up about 2.3 index points to Opus 5 xhigh while costing less than one third as much in this benchmark.
For organizations operating coding agents at meaningful volume, this may be the most commercially important point in the update. Frontier performance matters, but a modest quality difference can be overwhelmed by a large cost difference when thousands of tasks are routed through the same setting.
The correct decision is workload-specific. Use the higher-cost route where failure or rework is expensive. Use the value route where tests, review, and retry policies make the remaining quality gap manageable.
Muse Spark strengthens the middle of the curve
Meta's Muse Spark 1.1 xhigh reaches 53.54 at USD 1.43 per task with 34.0k output tokens in the Opencode harness.
Spark is not a substitute for Opus 5 on the hardest tasks. Its relevance is economic: it gives engineering teams another measured option between low-cost routine models and high-cost frontier agents.
Meta lists Muse Spark 1.1 in public preview with a 1 million token context window and Model API pricing of USD 1.25 per million input tokens, USD 4.25 per million output tokens, and USD 0.15 per million cache-hit input tokens. Those prices make it worth testing for bounded implementation, repair, migration, and test-generation workloads where the acceptance criteria are objective.
Source: Meta's Muse Spark 1.1 announcement.
Qwen 3.8 is real, but the chart point is not ready
Alibaba Model Studio now lists `Qwen3.8-Max-Preview`. The hosted preview is available through Alibaba's Token Plan and is positioned for agentic work.
What is still missing is more important for this analysis:
- No public Artificial Analysis Coding Agent Index v1.3 run.
- No stable pay-as-you-go token pricing in Alibaba's standard model-pricing table.
- No measured output-token footprint from the same coding harness.
- No fixed production model card or independently reproducible benchmark package.
Without those inputs, assigning Qwen 3.8 a cost and intelligence coordinate would create false precision. We have added it to the Friday update and source inventory, but not to the master scatter plot.
The right action is to evaluate the preview on a frozen internal workload, record the exact model identifier and effort setting, run each task multiple times, and retain tests, cost, tokens, latency, and human-review results. It should move onto the public plot when comparable evidence exists.
Source: Alibaba Model Studio.
The Friday routing decision
The current evidence supports a practical ladder:
- Route routine, well-tested work to Luna high or another efficient point that clears the acceptance threshold.
- Test Muse Spark for bounded agentic work where cost and speed matter more than frontier capability.
- Use Grok high as a strong upper-tier value route.
- Compare Sol max and Opus 5 xhigh on the organization's hardest representative tasks before standardizing either one.
- Treat Qwen 3.8 as a preview evaluation target, not a benchmark winner.
The chart is a starting point, not a procurement answer. Input caching, retries, tool calls, latency, failure rates, review time, and repository-specific correctness can change the real economics.
Explore all measured and modeled settings in the Aeon Model Economics Lab, then download the CSV or JSON dataset. For a workload-specific model benchmark and routing analysis, contact info@airiskmanagement.ca.