SN

私人新聞日報Private News Daily

Private use only · Michael SoMichael So

本頁為 Michael So 私人自用,非公開發佈,亦非任何機構之產品。 This page is for the personal use of Michael So only. It is not a public release and is not a product of any organisation.

AI 科技AI & Tech

谷歌發布 Gemini 4 Argon 旗艦模型,分階段開放、主打編程與網安Google releases Gemini 4 Argon flagship model, rolled out in phases, focusing on programming and cybersecurity

Google 官方部落格;VentureBeat;MarkTechPost ·2026-09-30

谷歌旗下人工智慧部門九月三十日發布新一代旗艦模型 Gemini 4 Argon,主打軟體工程、企業知識工作同網路安全防禦。輸出長度由六萬四千詞元升至一百萬詞元;初始定價每百萬輸入詞元兩美元、輸出十美元。首輪只開放畀可信網路防禦人員,其後才向開發者同消費者推出。Google's artificial intelligence division released its next-generation flagship model, Gemini 4 Argon, on September 30, focusing on software engineering, enterprise knowledge work, and cybersecurity defense. The output length has increased from 64,000 tokens to one million tokens; the initial pricing is two dollars per million input tokens and ten dollars per million output tokens. The first round is only open to trusted cybersecurity personnel, and it will be made available to developers and consumers later.

Last updated: October 2026

What is Gemini 4 Argon, and can you use it yet?

Gemini 4 Argon is Google DeepMind's new frontier model, announced on September 30, 2026. Today it is limited to trusted cyber defenders through the Fairwind Program, with paid API customers and Google AI Ultra subscribers next and no date given, according to Google (2026).

This guide covers access, pricing, vendor-reported and independent benchmarks, and the security questions a CISO should ask before approving the model. Every figure is dated and linked to the page where it appears.

TL;DR: Key Takeaways

- Access is narrow. Gemini 4 Argon reached only Fairwind cyber defenders at launch, and DataCamp (2026) found no published API model ID on September 30.

- Price undercuts the rivals at launch. Google lists $2 per million input tokens and $10 per million output tokens, rising to $4 and $20, with a 95% discount on cached input (Google, 2026).

- Output is the headline spec. The output limit is 1 million tokens, up from 64,000, and Google has not stated an input context window (MarkTechPost, 2026).

- Benchmarks are mixed. Google reports 77.9% on DeepSWE v1.1 against 74.2% for Claude Opus 5.5, yet Opus 5.5 leads Terminal-Bench 4.0 at 66.4% versus 57.4%, and no third party had reproduced Google's table at launch (DataCamp, 2026).

- Independent boards disagree. Artificial Analysis scores Argon 52.6 against 57.6 for Opus 5.5, while Vals ranks it first of 41 at 68.9% (Trending Topics, 2026; Vals AI, 2026).

- Security claims need testing. Google says Argon leads Gray Swan's indirect prompt injection benchmark but published no score, so enterprises should run their own tests.

At a glance: Argon against Opus 5.5 and GPT-6 Astra

Sources: Google, DataCamp, VentureBeat, Trending Topics, MarkTechPost (output limits), NeuralTrust. Retrieved October 1, 2026.

What is Argon and who can use it today?

Gemini 4 Argon is a Google DeepMind frontier model announced on September 30, 2026 and released first to trusted cyber defenders through the Fairwind Program. Google says paid API customers and Google AI Ultra subscribers follow "as soon as possible", while the model also takes part in the U.S. government's voluntary pre-release review (Google, 2026).

Where Argon sits in the Gemini line

Google's last flagship series was Gemini 3 in November 2025, according to VentureBeat (2026). Fello AI (2026) adds that Argon is the first Gemini with a codename instead of the Pro, Flash and Flash-Lite tiers, and that Google has not said what other Gemini 4 models will be called.

Who has access right now

DataCamp (2026) reports that on launch day Argon was absent from OpenRouter, Vertex AI, Gemini CLI, Cursor and GitHub Copilot. Engineering teams cannot move production traffic to it yet. Security teams are the exception: Google says Wiz used Argon through its Scan for Good initiative and found a critical vulnerability in healthcare software that earlier frontier models had missed.

Argon pricing: what does it cost against Opus 5.5 and GPT-6 Astra?

During the introductory period Gemini 4 Argon costs $2 per million input tokens and $10 per million output tokens. Standard rates are $4 and $20, with a 95% discount on cached input. Google has not said how long the introductory price lasts. At standard rates Argon matches Claude Opus 5.5.

Sources: Fello AI (cached prices), VentureBeat (Opus 5.5 and Astra prices), NeuralTrust. The last column is our arithmetic from the listed rates.

Sticker price is only part of the bill. According to Trending Topics (2026), Artificial Analysis measured a cost per average task of $1.99 for Argon, $3.26 for GPT-6 Astra and $5.98 for Claude Opus 5.5. The same evaluation shows Argon used 110 million tokens to run the index against a median of 82 million, so verbosity eats into the saving. Artificial Analysis also notes that Argon's price is introductory. DataCamp (2026) adds that Google has not said how reasoning tokens are billed.

Gemini 4 Argon benchmarks: what Google reports

Google reports Gemini 4 Argon ahead on DeepSWE v1.1 at 77.9%, the Vals Index at 68.9%, AutomationBench at 51.3% and Harvey's Legal Agent benchmark, and behind on FrontierSWE v2 and Terminal-Bench 4.0. These are vendor-reported figures, and DataCamp (2026) notes that no third party had reproduced them at launch.

Source: DataCamp (2026), compiling Google's announcement. Retrieved October 1, 2026.

The pattern is specific. Argon leads on long-horizon software tasks, legal and finance work, long-context retrieval and business automation. MarkTechPost (2026) reports 51.3% on AutomationBench against 42.5% for Opus 5.5. It trails where agents work in a terminal or across a large codebase: a 10.5-point gap to GPT-6 Astra on FrontierSWE v2 and a 9-point gap to Opus 5.5 on Terminal-Bench 4.0. Google also reports 91.7% on LVBench for long video understanding.

How we compared

We did not run these benchmarks. We use vendor-reported numbers where they are the only source and label them as such. Where an independent board exists we show it beside the vendor figure. Prices come from the vendors or from a named outlet, and every table carries its retrieval date.

What independent evaluations say about Argon

Independent evaluations place Gemini 4 Argon at or near the top on some boards and behind Claude Opus 5.5 on others. Artificial Analysis scores it 52.6 against 57.6 for Opus 5.5, while Vals ranks it first of 41 models on its index at 68.9%. The honest summary is a strong, uneven model.

Sources: Trending Topics (2026) for Artificial Analysis, Vals AI (2026). Artificial Analysis tested only Argon's High setting. Retrieved October 1, 2026.

The Decoder (2026) reports a hallucination rate of 15% for Argon against 51% for GPT-6 Astra, but factual accuracy of 50% against 63%, and a first place on the Arena.ai Text Arena at 1,525 points. It concludes that Google closes the gap without taking a clear lead. Axios (2026) adds that Bloomberg reported some Google staff found performance lacking in internal testing, a claim Google disputed.

Context window and 1M-token output limit

Google states a 1 million token output limit for Gemini 4 Argon, up from 64,000 on earlier Gemini models, and does not state an input context window. Some outlets report a 1 million token input window, but we could not match that figure to a Google source, so treat it as unconfirmed.

The output figure is the real differentiator. MarkTechPost (2026) reports that Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra each cap a single response at 128,000 tokens. Google cites internal migrations of C and C++ code to Rust at up to 800,000 lines, including the Fuchsia Zircon kernel, as the target use case.

A response of that size has a cost and a risk. At $10 per million output tokens a maximal answer costs up to $10 at introductory pricing and $20 at standard pricing. More important, an agent that writes 800,000 lines in one pass produces a diff no human can review line by line. The output limit raises the stakes of every guardrail around the agent.

Gemini 4 Argon security: what enterprises need to know

Google reports that Gemini 4 Argon leads Gray Swan's indirect prompt injection benchmark and ships with chain-of-thought and action monitors that can stop execution. It published no injection score, and DataCamp (2026) reports that the Fairwind cohort received the model without cyber safeguards. Model-level defenses help, but they do not replace controls around the agent.

What Google says it built in

Google lists refusal training for harmful requests, safeguards for cyber and CBRN misuse under its Frontier Safety Framework, internal and external red teaming, and activation monitoring to detect misuse. It adds automated red teaming and adversarial training against prompt injection, plus monitors that watch reasoning and actions and can halt a run. DataCamp (2026) notes Google kept monitoring findings out of training so the model does not learn to evade them.

……(原文過長,此處截斷)

Last updated: October 2026

What is Gemini 4 Argon, and can you use it yet?

Gemini 4 Argon is Google DeepMind's new frontier model, announced on September 30, 2026. Today it is limited to trusted cyber defenders through the Fairwind Program, with paid API customers and Google AI Ultra subscribers next and no date given, according to Google (2026).

This guide covers access, pricing, vendor-reported and independent benchmarks, and the security questions a CISO should ask before approving the model. Every figure is dated and linked to the page where it appears.

TL;DR: Key Takeaways

- Access is narrow. Gemini 4 Argon reached only Fairwind cyber defenders at launch, and DataCamp (2026) found no published API model ID on September 30.

- The price undercuts competitors at launch. Google charges $2 per million input tokens and $10 per million output tokens, increasing to $4 and $20, with a 95% discount on cached inputs (Google, 2026).

- The output is the headline specification. The output limit is 1 million tokens, up from 64,000, and Google has not indicated an input context window (MarkTechPost, 2026).

- Benchmarks are mixed. Google reports 77.9% on DeepSWE v1.1 compared to 74.2% for Claude Opus 5.5, yet Opus 5.5 leads Terminal-Bench 4.0 at 66.4% versus 57.4%, and no third party had reproduced Google's table at launch (DataCamp, 2026).

- Independent boards disagree. Artificial Analysis scores Argon 52.6 compared with 57.6 for Opus 5.5, while Vals ranks it first out of 41 at 68.9% (Trending Topics, 2026; Vals AI, 2026).

- Security claims need testing. Google says Argon leads Gray Swan's indirect prompt injection benchmark but published no score, so enterprises should run their own tests.

At a glance: Argon versus Opus 5.5 and GPT-6 Astra

Sources: Google, DataCamp, VentureBeat, Trending Topics, MarkTechPost (output limits), NeuralTrust. Retrieved October 1, 2026.

What is Argon and who can use it today?

Gemini 4 Argon is a Google DeepMind frontier model announced on September 30, 2026, and initially released to trusted cyber defenders through the Fairwind Program. Google says that paid API customers and Google AI Ultra subscribers will follow "as soon as possible," while the model also participates in the U.S. government's voluntary pre-release review (Google, 2026).

Where Argon sits in the Gemini line

According to VentureBeat (2026), Google's last flagship series was Gemini 3 in November 2025. Fello AI (2026) adds that Argon is the first Gemini with a codename instead of the Pro, Flash, and Flash-Lite tiers, and that Google has not said what the other Gemini 4 models will be called.

Who has access right now

DataCamp (2026) reports that on launch day, Argon was absent from OpenRouter, Vertex AI, Gemini CLI, Cursor, and GitHub Copilot. Engineering teams cannot move production traffic to it yet. Security teams are the exception: Google says Wiz used Argon through its Scan for Good initiative and found a critical vulnerability in healthcare software that earlier frontier models had missed.

Argon pricing: what does it cost compared to Opus 5.5 and GPT-6 Astra?

在介绍期,Gemini 4 Argon 每百万输入令牌收费 2 美元,每百万输出令牌收费 10 美元。标准费率为 4 美元和 20 美元,缓存输入可享受 95% 的折扣。Google 尚未说明介绍价格将持续多久。在标准费率下,Argon 与 Claude Opus 5.5 相匹配。

Sources: Fello AI (cached prices), VentureBeat (Opus 5.5 and Astra prices), NeuralTrust. The last column is our calculation based on the listed rates.

Sticker price is only part of the bill. According to Trending Topics (2026), Artificial Analysis measured a cost per average task of $1.99 for Argon, $3.26 for GPT-6 Astra and $5.98 for Claude Opus 5.5. The same evaluation shows Argon used 110 million tokens to run the index against a median of 82 million, so verbosity eats into the saving. Artificial Analysis also notes that Argon's price is introductory. DataCamp (2026) adds that Google has not said how reasoning tokens are billed.

Gemini 4 Argon benchmarks: what Google reports

Google reports Gemini 4 Argon ahead on DeepSWE v1.1 at 77.9%, the Vals Index at 68.9%, AutomationBench at 51.3%, and Harvey's Legal Agent benchmark, and behind on FrontierSWE v2 and Terminal-Bench 4.0. These are vendor-reported figures, and DataCamp (2026) notes that no third party had reproduced them at launch.

Source: DataCamp (2026), compiling Google's announcement. Retrieved October 1, 2026.

The pattern is specific. Argon leads in long-horizon software tasks, legal and finance work, long-context retrieval, and business automation. MarkTechPost (2026) reports 51.3% on AutomationBench, compared to 42.5% for Opus 5.5. It falls behind where agents operate in a terminal or across a large codebase: a 10.5-point gap with GPT-6 Astra on FrontierSWE v2 and a 9-point gap with Opus 5.5 on Terminal-Bench 4.0. Google also reports 91.7% on LVBench for long video understanding.

How we compared

我们没有运行这些基准测试。我们使用供应商报告的数字,当它们是唯一来源时,并标注为此。当存在独立机构时,我们将其显示在供应商数字旁边。价格来自供应商或指定的销售点,每张表格都有其获取日期。

What independent evaluations say about Argon

独立评价将 Gemini 4 Argon 在一些排行榜上列为第一或接近第一,而在其他排行榜上则落后于 Claude Opus 5.5。人工分析给它的评分是 52.6,而 Opus 5.5 为 57.6,而 Vals 则将其在 41 个模型的指数中排名第一,得分为 68.9%。诚实的总结是这是一个强大但不均衡的模型。

Sources: Trending Topics (2026) for Artificial Analysis, Vals AI (2026). Artificial Analysis tested only Argon's High setting. Retrieved October 1, 2026.

《解码器》(2026) 报告了 Argon 的幻觉率为 15%,而 GPT-6 Astra 的幻觉率为 51%,但事实准确率为 50%,而 GPT-6 Astra 为 63%,并且在 Arena.ai 文本竞技场上以 1,525 分名列第一。它的结论是,谷歌缩小了差距,但没有取得明显领先。《Axios》(2026) 补充说,彭博社报道,一些谷歌员工在内部测试中发现性能不足,而谷歌对此提出异议。

Context window and 1M-token output limit

Google states a 1 million token output limit for Gemini 4 Argon, up from 64,000 on earlier Gemini models, and does not state an input context window. Some outlets report a 1 million token input window, but we could not match that figure to a Google source, so treat it as unconfirmed.

The output figure is the real differentiator. MarkTechPost (2026) reports that Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra each limit a single response to 128,000 tokens. Google cites internal migrations of C and C++ code to Rust of up to 800,000 lines, including the Fuchsia Zircon kernel, as the target use case.

A response of that size has a cost and a risk. At $10 per million output tokens, a maximal answer costs up to $10 at introductory pricing and $20 at standard pricing. More importantly, an agent that writes 800,000 lines in one pass produces a diff that no human can review line by line. The output limit increases the stakes of every safeguard around the agent.

Gemini 4 Argon security: what enterprises need to know

Google reports that Gemini 4 Argon leads Gray Swan's indirect prompt injection benchmark and comes with chain-of-thought and action monitors that can stop execution. It did not publish an injection score, and DataCamp (2026) reports that the Fairwind cohort received the model without cyber safeguards. Model-level defenses help, but they do not replace controls around the agent.

What Google says it built in

Google lists refusal training for harmful requests, safeguards for cyber and CBRN misuse under its Frontier Safety Framework, internal and external red teaming, and activation monitoring to detect misuse. It adds automated red teaming and adversarial training against prompt injection, plus monitors that watch reasoning and actions and can halt a run. DataCamp (2026) notes Google kept monitoring findings out of training so the model does not learn to evade them.

……(The original text is too long, truncated here)

原文出處:Source: Google 官方部落格;VentureBeat;MarkTechPost ↗