
Google’s Gemini 4 Argon is level with OpenAI’s GPT-6 Astra in early independent tests. It scored 53 on the Artificial Analysis Intelligence Index, the same as GPT-6 Astra and one point ahead of GPT-6.1 Sol, Artificial Analysis said. Anthropic’s Claude Opus 5.5 still leads the index, with 58.
The score is 23 points above Gemini 3.1 Pro Preview, Google’s last model above the Flash class. The benchmarking firm said it puts Google back among the top three AI labs. Google unveiled Argon on Wednesday and has so far given it only to a small group of cyber defenders.
Fewer hallucinations, more tokens
Argon had a hallucination rate of 15% on AA-Omniscience, a test of factual knowledge. That is the lowest of any model scoring 45 or more on the index. GPT-6 Astra came in at 51% and GPT-6.1 Sol at 54%. Argon more often admits it does not know an answer instead of guessing, the firm said. It also got fewer answers right: 50%, against 63% for GPT-6 Astra.
Argon ranked first on AutomationBench-AA, a test of business software workflows, with 78%. That is seven points ahead of Claude Sonnet 5.5. On Terminal Bench 4, a coding test, it scored 57%. Claude Sonnet 5.5 (64%), Claude Opus 5.5 (60%) and GPT-6 Astra (59%) all scored higher.
Each index task cost $1.99 with Argon at its launch discount, 60% of GPT-6 Astra’s $3.26. The saving comes from lower token prices, not from using fewer tokens. Argon used about 62,000 output tokens per task, against 27,000 for GPT-6 Astra. It still costs 2.7 times as much per task as GPT-6.1 Sol.
Google is charging half price for at least a month: $2 per million input tokens and $10 per million output tokens. At the standard $4 and $20, each task rises to $3.98, about 1.2 times GPT-6 Astra, according to Artificial Analysis. Google has not said when the discount ends.
First on the Arena text ranking
Argon also leads Arena’s Text Arena, where people vote between model answers. It has 1,525 points, 20 ahead of Claude Opus 4.6, The Decoder reported. It placed eighth in Arena’s WebDev ranking. It is also the first Gemini model to top the Vals Index, at 68.9%.
Doubts inside Google
Some Google staff are less sure of the model. Its coding abilities are uneven, and it is not particularly good at front-end design, one person said. Front-end design shapes how apps and websites look. Julia Love and Davey Alba reported this for Bloomberg, citing people familiar with the internal evaluations.
Argon is also a very large model, according to Bloomberg, and large models are typically expensive to run. Two people said it appears affected by “benchmaxxing”, when engineers chase a test score over a product that does the job well.
Some employees believe Anthropic’s Fable and OpenAI’s Astra models are improving faster than Gemini. Others believe Gemini 4 has caught up. Google told Bloomberg it would be inaccurate to say the model underperforms in areas such as coding. It pointed to remarks by Koray Kavukcuoglu, who runs Google DeepMind, at a conference hosted by The Information.
“In my mind, it’s a certainty that we are always gonna be at the frontier,” Kavukcuoglu said.
Kavukcuoglu had said Google wanted Gemini 4 out much earlier than year-end. Paid API customers and Google AI Ultra subscribers are next in line for Argon. Google has not given a date.