AI Research - 6 min read

GPT-6.1 Sol vs Gemini 4 Argon: Coding Agent Costs and Access

Compare Sol 6.1 effort levels and Argon coding-agent costs, with independent results, introductory pricing and restricted-access caveats.

GPT-6.1 Sol is a compelling lower-cost coding pilot. Gemini 4 Argon offers a higher independent coding-agent point estimate, but at greater task cost and with restricted initial access. Neither launch makes maximum reasoning effort the automatic default. The useful buying question is cost per accepted outcome, including review and failure recovery.

We have added both releases to the complete default Model Economics Lab, including source-linked diamonds, five observed Sol effort levels and downloadable data. This analysis uses sources checked on October 1, 2026. It is not an Aeon-run head-to-head benchmark.

The independent coding-agent comparison

Artificial Analysis Coding Agent Index v1.5 combines DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA with equal weights. These are native agent systems: Codex for Sol and Antigravity CLI for Argon, not the same model-only harness.

System and effortIndexUSD per evaluated taskOutput tokens per task
GPT-6.1 Sol Low57.22$0.5014.7k
GPT-6.1 Sol Medium61.41$0.7022.4k
GPT-6.1 Sol High60.15$0.8928.7k
GPT-6.1 Sol XHigh62.91$1.0435.4k
GPT-6.1 Sol Max60.15$1.5561.7k
Gemini 4 Argon High63.76$5.84112.1k

The Sol XHigh point is our opening configuration; Medium is a useful budget pilot. Max uses more tokens and costs more without a higher observed score. Argon's 0.85-point lead over XHigh costs roughly 5.6 times as much per evaluated task at launch pricing. Small differences are not established statistical superiority, and the agent harnesses differ. The composite is not a pass percentage or cost per successful task.

The snapshot records no safety fallbacks for either new system. Sol XHigh has 13 hard stops among 909 attempts; Argon has none among 905. Each has one flagged Terminal-Bench attempt in the published reward-hacking audit. This is benchmark telemetry, not evidence of universal safety or a reason to remove deployment controls.

Sol 6.1: price the actual route

OpenAI's model documentation lists standard API prices of $2 input and $10 output per million tokens, a 1,050,000-token context window and a 128,000-token maximum output. Cached input is $0.10 per million; cache writes have separate pricing. Requests above 272,000 input tokens have higher rates. Service tiers also change billing. Subscription allowances are not these API task costs.

OpenAI's launch evaluation says Sol 6.1 matches Astra on DeepSWE v1.1 at roughly one-fifth of the task cost. This is a vendor-reported comparison, not the independent composite. Do not splice a launch score into another run's cost. The practical opportunity is an economical route for bounded delegated work, with escalation when tests fail or requirements are ambiguous. OpenAI lists availability in the API, Codex and ChatGPT Work, but not yet Chat.

Argon: frontier capability with two commercial caveats

Google's September 30 announcement reports 77.9% DeepSWE v1.1 and 51.3% AutomationBench. These are vendor-reported results, separate from the independent index. Google also describes an expanded one-million-token output limit, not merely a context-window claim. More generation headroom can support long work, but is not inherently a productivity gain.

Access initially goes to trusted testers and Fairwind cyber defenders. Broad availability is planned, not established by the announcement. The introductory API rate is $2 input / $10 output per million tokens; the footnote lists $4 / $20 after the introductory period. We retain the published $5.84 task observation with that caveat rather than inventing a post-promotion task coordinate.

Google describes prompt-injection defenses, action monitoring and hardened evaluation environments. Those controls do not remove the buyer's responsibility for authorized targets, sandboxing, data handling and approval gates.

How to run a useful buying pilot

Use the same repository tasks, tool permissions and acceptance tests across the systems you can actually access. Include routine changes, ambiguous debugging and one long-horizon task class. Keep sensitive data out unless the approved account and contract cover it.

Record accepted changes, failed attempts, wall-clock time, API charges, reviewer minutes and escaped defects. Define success before running the pilot. A cheaper task is not cheaper work if review or rework erases the saving; a more capable system is not valuable if its access or controls do not fit the workflow.

Start Sol at Medium for bounded work and test XHigh where quality requirements justify it. Evaluate Argon for genuinely difficult workflows when authorized access exists. Keep a fallback and a human gate for consequential actions. Do not automate routing on a sub-one-point benchmark difference.

What changed in the Lab

The main view now contains 16 current and maintained shortlist models. All/history retains 21 families and 38 operating points. The new releases appear as diamonds with native metric, harness, access and price notes. Historical v1.3 circles, independent mini-swe-agent DeepSWE rows and CursorBench history remain separate evidence cohorts; their dates are retained rather than relabeled as fresh measurements.

Explore the interactive chart and full dataset. For an organization-specific workload and control review, see Aeon's AI Control ROI Assessment. Choose models by accepted work and governed operating cost, not launch headlines alone.