Grok 4.7 deserves a fresh look. Claude Opus 5.5 strengthens the case for a premium subscription. GPT-6 Sol and Luna make selective delegation more affordable. Together, these releases make one-model-for-everything purchasing harder to justify.
We have updated the Aeon Model Economics Lab with all four. The practical question is not which logo wins. It is how much accepted work your organization gets from its budget, subscription capacity and review time.
Grok 4.7: a credible catch-up story
SpaceXAI released Grok 4.7 on September 21. Its published CursorBench 4.0 result rises from 40.4% for Grok 4.6 High to 46.3% for Grok 4.7 XHigh. Those are different effort settings, so the gain should not be treated as a controlled model-only experiment. Official announcement.
There is independent support for the direction. In Artificial Analysis's current Coding Agent Index v1.5, Grok Build with Grok 4.7 XHigh scores 56.27 versus 46.97 for 4.6 XHigh. Its measured API task cost also rises, from $3.57 to $8.82, alongside higher output-token use. Catching up is not the same as becoming cheaper to finish every task. Independent coding-agent results.
Our assessment: Grok belongs on the shortlist for a new coding pilot, particularly where teams want a credible alternative provider. Test it against real repository tasks, not brand reputation. Retain human review and compare failed attempts as well as successful patches.
Opus 5.5: more useful subscription capacity
Anthropic's September 22 announcement combines two distinct changes. It estimates approximately 40% lower cost than Opus 5 on typical workloads at default settings. Separately, it announces higher five-hour limits on Pro, Max, Team and seat-based Enterprise plans, plus a rate-limit reset subscribers can save. Anthropic's release announcement.
That is a meaningful reason to reassess subscription value. A team repeatedly interrupted by a usage ceiling may benefit from more sustained working sessions. But neither the lower API rate nor the announcement establishes a universal percentage increase in included messages, weekly capacity or completed projects. We have not measured a subscription-throughput multiplier.
There is also a demanding-workload caveat. At maximum effort, the current independent Claude Code comparison puts Opus 5.5 at 65.99 and $13.04 per task, versus Opus 5 at 59.73 and $10.79. The new result includes safety recovery and fallback attempts. This is stronger benchmark performance, not evidence that every task is 40% cheaper. Claude Code and Codex comparison.
Our view: Opus 5.5 has an attractive subscription-value story, especially for demanding daily work. Measure accepted outputs per billing period and time lost to limits before changing seat counts.
GPT-6 Sol and Luna: lower-cost delegation candidates
OpenAI lists GPT-6 Sol at $2 input and $10 output per million tokens, with cached input at $0.20. GPT-6 Luna is $0.10 input, $0.50 output and $0.01 cached input. Both support adjustable reasoning. These are base API rates, not subscription allowances. Sol documentation, Luna documentation.
Sol is a candidate for bounded implementation and debugging. Luna is worth testing on classification, extraction, simple transformations and other tightly specified work. Those are proposed workload assignments, not claims that either model passes your acceptance tests without supervision.
The four new plot coordinates
All four rows below use Artificial Analysis Coding Agent Index v1.5, accessed September 25. They share an evaluation suite, but run in different native agent harnesses. Cost means average API cost per evaluated task, including unsuccessful attempts, not cost per successful business outcome or subscription debit. Marker area uses output tokens, not total input-plus-cache-plus-output tokens.
| Model and harness | Effort | Coding index | API cost/task | Output tokens/task |
|---|---|---|---|---|
| Grok 4.7, Grok Build | XHigh | 56.27 | $8.82 | 123.3k |
| Claude Opus 5.5, Claude Code | Max | 65.99 | $13.04 | 333.2k |
| GPT-6 Sol, Codex | Max | 56.66 | $2.99 | 62.2k |
| GPT-6 Luna, Codex | Max | 41.07 | $0.18 | 98.3k |
Source: Artificial Analysis coding-agent measurements and methodology. Values are rounded. Opus 5.5's published run records 78 fallback and two continued attempts among 909 retained attempts; it is not a no-fallback model isolation.
The default master plot now shows 14 current models. Grok 4.7, Opus 5.5 and GPT-6 Sol replace their predecessors in that view; GPT-6 Luna joins them. Earlier observations remain in All / history. The four additions are diamonds because the retained circles use an older benchmark version. Do not interpret vertical placement across versions as a controlled ranking. The four v1.5 rows can be compared with one another as tested systems.
How companies should use this update
Run a small, repeatable pilot before expanding spend:
- Select representative tasks with clear acceptance tests and permitted data access.
- Compare a premium lead model with a cheaper model on the same bounded work. Include Grok where provider diversification matters.
- Record accepted outputs, corrections, review minutes, elapsed time and failures. Keep API dollars and subscription-limit consumption separate.
- Escalate difficult work deliberately. Do not make every task maximum effort by default.
- Retest when the model, harness or pricing changes.
For a fixed subscription, divide the seat cost and any overage by accepted outputs during the billing period. For APIs, include retries, tools and supervision. Neither metric alone measures confidentiality, security or accountability, which need their own acceptance criteria.
Aeon helps organizations turn these choices into working AI deployments: task routing, private AI where appropriate, security controls and governance evidence. Discuss an AI implementation and model-selection pilot, or explore the updated interactive plot.
Research date: September 25, 2026. USD API prices exclude taxes and additional tool charges. Cache writes, long-context premiums, processing modes and regional pricing can change the bill. Published benchmarks are not Aeon client outcome claims; subscription value depends on the plan and actual usage.