AI Research - 9 min read

GPT-6 Astra and Gemini 3.8 Flash Tie. Their Economics Do Not.

Astra, Fable 5.1, Gemini 3.8 Flash, GLM-5.3, Muse Spark 1.3, and Qwen3.8-Max-0902 show why coding score, completed-task cost, and token footprint need separate evidence.

Six current model releases changed the coding-agent economics map in one week: GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, GLM-5.3, Muse Spark 1.3, and Qwen3.8-Max-0902.

The headline result is not a universal winner. It is a widening gap between benchmark capability, completed-task cost, and workload footprint.

In the current independent DeepSWE v1.1 run, GPT-6 Astra and Gemini 3.8 Flash both reach 74 percent. Gemini averages USD 2.36 per task, while Astra averages USD 6.52. Astra uses about 30k output tokens per task versus Gemini's 143k. The quality result is a tie within the published point estimates, but the operating profiles are materially different.

We updated the Aeon Model Economics Lab to make those tradeoffs visible. The approved current view now contains 13 models. GPT-5.6 Sol remains; GPT-5.6 Terra, GPT-5.6 Luna, GLM-5.2, Fable 5, Gemini 3.7 Flash, Muse Spark 1.1 and 1.2, and the undated Qwen3.8 Max entry have been removed.

The six new operating points

The master plot combines multiple evidence families for discovery, so each point retains its native benchmark and token definition. Diamonds identify points that are newer-suite, cross-suite, or proxied rather than directly comparable with the retained Artificial Analysis v1.3 circles.

Model and settingNative coding scoreCost per taskOutput tokensEvidence basis
GPT-6 Astra xhigh74%USD 6.5230kIndependent DeepSWE v1.1
Gemini 3.8 Flash high74%USD 2.36143kIndependent DeepSWE v1.1
Claude Fable 5.1 max70.4USD 9.18101.5kArtificial Analysis Coding Agent Index v1.4, with fallback
GLM-5.3 max69%USD 3.9980kIndependent DeepSWE v1.1
Muse Spark 1.3 xhigh64.2USD 1.7268.2kArtificial Analysis Coding Agent Index v1.4
Qwen3.8-Max-0902 xhigh57% proxyUSD 3.73 proxy95k proxyPredecessor Qwen3.8 Max DeepSWE run

These numbers are benchmark operating points, not a single ranking. DeepSWE and the Artificial Analysis Coding Agent Index use different task mixtures, agent implementations, scoring, and cost accounting. Fable 5.1 and Muse Spark 1.3 also use the newer Artificial Analysis v1.4 methodology, while the filled historical circles in the lab use v1.3.

Astra and Gemini tie on DeepSWE, then diverge

The independent DeepSWE leaderboard reports 74 percent for GPT-6 Astra at xhigh and 74 percent for Gemini 3.8 Flash at high. The published uncertainty ranges overlap, so presenting one as the coding winner would overstate the evidence.

Their economics are less ambiguous.

Gemini's USD 2.36 average task cost is about 64 percent below Astra's USD 6.52. Astra, however, produces roughly one fifth as many output tokens. That smaller output footprint may matter where latency, review burden, context transfer, or downstream processing is expensive.

OpenAI's release page confirms the GPT-6 Astra model and pricing. Google's Gemini 3.8 Flash model card documents the model and its intended operating profile. The procurement question is therefore not simply which score is higher. It is whether the deployed workload values lower completed-task cost, smaller output volume, latency, or a particular harness integration.

Full GLM-5.3 and Flash define a routing ladder

GLM-5.3 and GLM-5.3-Flash belong in the same buying decision, but not in the same role.

DeepSWE reports 69 percent, USD 3.99 per task, and 80k output tokens for full GLM-5.3. GLM-5.3-Flash reaches 63.4 percent at USD 0.24 per task with 73k output tokens. That is a 5.6-point score gap and a more than 16-fold task-cost gap in the same harness.

The practical policy is straightforward:

  1. Route bounded, high-volume work to Flash when it clears the acceptance threshold.
  2. Escalate complex dependency chains and stubborn failures to full GLM-5.3.
  3. Measure accepted artifacts, retries, and human review time rather than optimizing raw benchmark score alone.

Z.ai's GLM-5.3 product documentation positions the flagship for long-horizon engineering work. Its GLM-5.3-Flash release emphasizes efficiency and confirms that the earlier Ox Alpha preview was GLM-5.3-Flash.

Fable 5.1 and Muse Spark 1.3 use the newer AA suite

Claude Fable 5.1 records a 70.4 Artificial Analysis Coding Agent Index v1.4 score at USD 9.18 per task and 101.5k output tokens. The tested Claude Code configuration includes fallback, so this is a system result, not a pure single-model isolation. Anthropic's Fable page confirms the release, while the Artificial Analysis comparison provides the task-level economics.

Muse Spark 1.3 records 64.2 on the same v1.4 index at USD 1.72 per task and 68.2k output tokens in Muse Code. That is a stronger value signal than the previous Muse entries, but it still needs workload-specific validation. Artificial Analysis publishes both the Muse Spark 1.3 analysis and the coding-agent run.

The lab uses diamonds for both models because its older measured circles come from Artificial Analysis v1.3. A new benchmark version is useful evidence, but it is not identical evidence.

Qwen3.8-Max-0902 is real; its plotted economics are still a proxy

Alibaba Cloud confirms that `qwen3.8-max-0902`, also called `qwen3.8-max-2026-09-02`, is an upgraded snapshot with a one-million-token context window.

A complete same-harness coding-agent economics run for that dated snapshot was not public when this update was prepared. The lab therefore does not present the predecessor result as a measured 0902 score.

The temporary diamond carries the preceding Qwen3.8 Max DeepSWE coordinate: 57 percent, USD 3.73 per task, and 95k output tokens. Every field is labeled as a predecessor proxy. It will be replaced when the dated snapshot receives a complete public run.

This distinction matters. A release announcement proves availability and identity. It does not prove completed-task economics.

What model buyers should do now

The current data supports a routing experiment, not a blanket vendor decision.

  1. Freeze a representative repository task set and acceptance rubric.
  2. Test Gemini 3.8 Flash and GLM-5.3-Flash as lower-cost default routes.
  3. Test Astra, Fable 5.1, Opus 5, Sol, and full GLM-5.3 on the difficult tail.
  4. Include Muse Spark 1.3 where its Muse Code harness fits the workflow.
  5. Keep Qwen3.8-Max-0902 in evaluation status until its dated task economics are measured.
  6. Record accepted-task cost, retries, elapsed time, output tokens, and reviewer minutes.

The main lesson is stable even as model names change: buy completed work, not benchmark rank. A model that is slightly weaker but clears the acceptance threshold at one tenth the task cost can be the better production choice. A more expensive model can still be rational when it reduces failure risk, review burden, or repeated attempts.

Explore the maintained Model Economics Lab, download the CSV or JSON dataset, or contact Aeon to benchmark a routing policy against your own repositories.