Three new coding-model developments arrived with three different kinds of evidence.
DeepSeek V4 Flash 0731 has an unusually low measured cost on the Artificial Analysis Intelligence Index. Meta's Muse Code, powered by Muse Spark 1.2, has strong vendor-reported coding results. Qwen3.8-Max-Preview is available in Qwen Code, but still lacks a comparable public coding-agent economics run.
The easy response is to place all three on one chart. The correct response is to keep unlike evidence separate.
We updated the Aeon Model Economics Lab with two complementary views:
- The master plot remains based on Artificial Analysis Coding Agent Index v1.3. Cost is the horizontal axis, coding score is the vertical axis, and output-token volume is encoded in bubble area.
- A separate CursorBench 3.2 comparison shows score, completed-task cost, total tokens, and agent steps from the same Cursor harness.
That structure makes the page more useful for buyers. It shows where rankings agree, where they diverge, and where a new model is promising but not yet comparable.
DeepSeek V4 Flash 0731 changes the cost conversation
Artificial Analysis reports an Intelligence Index v4.1 score of 49.93 for DeepSeek V4 Flash 0731 at maximum reasoning effort. The same run reports an average general-task cost of about USD 0.027 and 46.3k output tokens per task.
That cost is commercially significant. It suggests that open-weight, mixture-of-experts systems can make high-volume reasoning workloads materially cheaper when the task quality threshold is moderate and the deployment can tolerate the model's operational profile.
It does not prove that DeepSeek V4 Flash is the best coding agent.
The 49.93 score comes from the general Artificial Analysis Intelligence Index v4.1, not the Coding Agent Index v1.3 used by our master coding plot. The task suite, score, and cost denominator are different. Plotting the point beside Opus 5, Sol, Grok, or Kimi would imply a comparison that the source data does not support.
Our decision is therefore deliberate: publish the measured general-economics evidence, label it clearly, and wait for a comparable coding-agent run before assigning a coding coordinate.
CursorBench gives model economics a second lens
CursorBench 3.2 evaluates ambiguous, multi-file coding tasks derived from real Cursor sessions. It reports score, average cost per task, total token use, and agent steps. That makes it valuable for model economics because the benchmark connects capability to actual workload consumption.
A selected frontier set tells a different story from the Artificial Analysis master plot:
| Model and setting | CursorBench score | Cost per task | Total tokens | Steps |
|---|---|---|---|---|
| Claude Fable 5 max | 70.5 | USD 17.32 | 103.5k | 72 |
| Claude Opus 5 max | 70.0 | USD 8.23 | 61.8k | 78 |
| GPT-5.6 Sol max | 67.2 | USD 5.69 | 28.3k | 48 |
| Grok 4.5 high | 66.7 | USD 1.51 | 19.5k | 33 |
| GPT-5.6 Terra max | 64.9 | USD 2.31 | 33.0k | 47 |
| GPT-5.6 Luna max | 61.1 | USD 0.39 | 88.0k | 61 |
| Kimi K3 max | 60.8 | USD 2.70 | 38.4k | 57 |
| GPT-5.5 high | 58.4 | USD 2.05 | 12.2k | 28 |
| Composer 2.5 | 56.1 | USD 0.44 | 14.3k | 33 |
| GLM-5.2 max | 55.0 | USD 1.76 | 35.9k | 58 |
This is not a replacement ranking. CursorBench and Artificial Analysis use different tasks, harnesses, scoring systems, and token definitions. Cursor also warns that Grok 4.5 may have benefited from an earlier Cursor codebase snapshot appearing in training data, with an unclear effect on the result.
The value is triangulation. When a model looks strong across independent harnesses, confidence rises. When rankings diverge, the buyer should test on its own repositories before committing volume.
Muse Code shows why the harness is part of the product
Meta introduced Muse Code and Muse Spark 1.2 as a co-designed agent and model system. Meta reports:
- 82.9 percent on Terminal-Bench 2.1.
- 59.3 percent on DeepSWE 1.1.
- 70.6 percent on Meta's internal coding benchmark.
These results are strong enough to justify evaluation. They also reinforce a strategic point: the model and the agent harness should be evaluated as one operating system.
Repository navigation, tool selection, context management, verification, and recovery behavior can change completed-task performance as much as raw model capability. Muse Code's co-training approach is an explicit bet that harness compatibility is a source of advantage.
The limitation is economic comparability. We do not yet have an independent common-harness cost-per-task and token point for Muse Spark 1.2 in Muse Code. Until that exists, the correct label is vendor benchmark evidence, not measured position on the master economics curve.
Qwen3.8 is available, but the chart point is still pending
Qwen Code confirms that `qwen3.8-max-preview` is available through its Token Plan.
Availability is useful evidence. It allows teams to run frozen internal tasks, record exact model identifiers, and compare acceptance rate, retries, latency, token consumption, and review burden.
It is not enough for a public economics coordinate. We still need a comparable coding-agent score, completed-task cost, and token measurement from the same run. Qwen3.8 remains on the evaluation watchlist until those inputs are public.
What buyers should do now
Model procurement should move from brand selection to workload routing.
- Define task classes and acceptance tests before choosing models.
- Use independent benchmarks to establish a shortlist, not a final answer.
- Run the shortlist in the actual harness that will be deployed.
- Measure completed-task cost, total tokens, latency, retries, and human review.
- Route routine work to the lowest-cost point that clears the quality threshold.
- Reserve frontier effort for ambiguity, large failure costs, and complex dependency chains.
The most important conclusion from this update is not that one model won. It is that model economics is inseparable from benchmark design, harness behavior, and workload policy.
DeepSeek V4 Flash 0731 may reset the cost floor for broad reasoning. Muse Code may raise the ceiling for co-designed model and harness systems. Qwen3.8 may become a compelling production option. Each claim needs the right evidence before it becomes a procurement decision.
Explore the maintained Model Economics Lab, download the dataset, or contact Aeon to benchmark models against your own workflows.