Skip to content
한국어
Stage · Developing2 outlets · 2 articles

Gemini 4 Argon leads 12 of 18 Google benchmarks; some staff say coding lags

Circulating claim

“Google's new Gemini 4 Argon model struggles in real-world use, despite its benchmark scores.”

Unverified

2 sources · verified by Newsmesis

Gemini 4 Argon leads 12 of 18 Google benchmarks; some staff say coding lags — Newsmesis editorial illustration
— Newsmesis AI illustration

What happened

Google's Gemini 4 Argon scored 77.9% on the DeepSWE v1.1 long-horizon software engineering benchmark, against 74.2% for Anthropic's Claude Opus 5.5 on Google's own chart. Bloomberg reported on Oct. 1 that some Google employees find the model does less well in daily work than on benchmarks, and that it struggles with certain coding tasks. Google told Bloomberg it would be inaccurate to say Gemini 4 is underperforming in areas such as coding. Argon is currently open only to vetted cybersecurity defenders, so outside testing of the real-world question is limited.

What we know

SiliconANGLEsiliconangle.com9to5Google9to5google.com
  • Per SiliconANGLE, Argon scored 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra a tenth of a point behind Opus, all on Google's chart.
  • Per SiliconANGLE, VentureBeat counted 12 of the 18 benchmarks in Google's charts as outright Argon wins; Opus 5.5 still leads Terminal-Bench 4.0, and Astra kept FrontierSWE v2.
  • Per 9to5Google, citing Bloomberg, Googlers said the model does less well when employees put it to work and struggles with certain coding tasks. Another employee described a large consensus that it is at the frontier.
  • Per SiliconANGLE, access is limited to Google's own teams and members of its Fairwind Program, which has signed up more than 650 organizations including CrowdStrike and Palo Alto Networks. Google has given no date for paying developers or AI Ultra subscribers.
  • Per SiliconANGLE, launch pricing is $2 per million input tokens and $10 per million output tokens, rising to Opus 5.5's $4 and $20 when the promotion ends. 9to5Google reports the same $4/$20 list price with half-price introductory rates.

What we don't know yet

  • How many employees hold each view. Bloomberg's account, as relayed by 9to5Google, quotes a few unnamed Googlers on both sides, which is not a measured split.
  • Whether the coding complaints concern specific task types or the model overall. Google disputes the underperformance framing, and no independent coding evaluation of Argon has been published.
  • Whether the benchmark leads hold in independent tests. The 18-benchmark comparison comes from Google's own charts, and the cyber-restricted version is not what most developers will get.
  • When broader access opens, and whether the model changes before then.

Why it matters

Developers choosing between Argon, Opus 5.5 and GPT-6 Astra for coding agents will have to rely on Google's own benchmark charts until independent tests are possible. Argon is already priced at Opus parity after the promotional period, so any gap between benchmark and real-world performance bears directly on its cost.

Claims

  • ConfirmedGemini 4 Argon scored 77.9% on DeepSWE v1.1, versus 74.2% for Claude Opus 5.5, on Google's own chart.
  • ConfirmedVentureBeat counted 12 of the 18 benchmarks in Google's charts as outright Argon wins.
  • ConfirmedBloomberg reported that Google employees found Gemini 4's real-world performance does not always match its benchmark scores.
  • UnconfirmedGemini 4 Argon struggles with certain coding tasks in real-world use.
  • ConfirmedGoogle said it would be inaccurate to say Gemini 4 is underperforming in areas such as coding.
  • ConfirmedArgon is available only to Google's internal teams and Fairwind Program members, which has more than 650 organizations.
  • ConfirmedArgon's launch price is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the promotion.
  • ConfirmedArgon's output limit is 1 million tokens, up from 64,000 on earlier Gemini models.

Related coverage

Why you can trust this story

65%

Source map · 2 outlets / 2 articles

How this credibility score is calculated

  • Source reliability
  • Corroboration
  • Primary source
  • Atom-verified claims
  • No contradiction
  • Settled
  • Claim attribution
  • AI disclosure

Atom-verified claims matched cited source text. Weights are fixed and explainable.

Human accountability

  • 2 independent origins / 2 sources
  • Drafted by the Newsmesis judgment agent (agent-cli); human approval required before publishing.

Verification ledger · 8

  • Claim: Gemini 4 Argon scored 77.9% on DeepSWE v1.1, versus 74.2% for Claude Opus 5.5, on Google's own chart.
    Status: ConfirmedVerification trailVerifiedChecked by verification engine

    At least one extracted atom matched cited source full text.

  • Claim: VentureBeat counted 12 of the 18 benchmarks in Google's charts as outright Argon wins.
    Status: ConfirmedVerification trailVerifiedChecked by verification engine

    At least one extracted atom matched cited source full text.

  • Claim: Bloomberg reported that Google employees found Gemini 4's real-world performance does not always match its benchmark scores.
    Status: ConfirmedVerification trailVerifiedChecked by verification engine

    At least one extracted atom matched cited source full text.

  • Claim: Gemini 4 Argon struggles with certain coding tasks in real-world use.
    Status: UnconfirmedVerification trailUnverifiedChecked by verification engine

    No atom-level source-text verification is available.

  • Claim: Google said it would be inaccurate to say Gemini 4 is underperforming in areas such as coding.
    Status: ConfirmedVerification trailVerifiedChecked by verification engine

    At least one extracted atom matched cited source full text.

  • Claim: Argon is available only to Google's internal teams and Fairwind Program members, which has more than 650 organizations.
    Status: ConfirmedVerification trailVerifiedChecked by verification engine

    At least one extracted atom matched cited source full text.

  • Claim: Argon's launch price is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the promotion.
    Status: ConfirmedVerification trailVerifiedChecked by verification engine

    At least one extracted atom matched cited source full text.

  • Claim: Argon's output limit is 1 million tokens, up from 64,000 on earlier Gemini models.
    Status: ConfirmedVerification trailVerifiedChecked by verification engine

    At least one extracted atom matched cited source full text.

Update log · 0

    AI accelerates. Humans approve.