According to Google, Argon demonstrates leadership across multiple coding and security benchmarks and identified a vulnerability that previous-generation models overlooked. However, Bloomberg reports that certain members of Google's own team harbor reservations about its effectiveness in practical applications.

The technology giant has introduced Gemini 4 Argon, marking the first major iteration of its core AI model since Gemini 3 debuted in November. Cybersecurity professionals will gain access to the system initially, with paid API users and Google AI Ultra subscribers following shortly thereafter, the company announced Wednesday.

Koray Kavukcuoglu, who oversees Google DeepMind's daily operations, described Argon as the firm's "next era of frontier intelligence" during the announcement. The model targets demanding, multi-step assignments spanning software development, legal services, financial analysis, and defensive cybersecurity operations.

Security professionals gain early access

The initial rollout reaches participants in Google's Fairwind Program, which brings together "trusted cyber defenders." Additionally, Google is participating in a voluntary United States government initiative for early-stage model evaluation. Broader availability will follow after additional validation phases.

These early adopters will receive Argon without its standard safety restrictions, enabling them to leverage its complete capabilities, Google explained. Security outfit Wiz has already begun deploying it through its Scan for Good program. The system uncovered a serious security gap in healthcare infrastructure deployed across hospital networks globally—a flaw that competing frontier-class models had failed to detect, Wiz reported.

Tulsee Doshi, leading Gemini product development at Google DeepMind, shared his perspective with Axios:

Argon is a well-rounded model that has frontier capabilities across several domains.

Tulsee Doshi, head of Gemini products at Google DeepMind

Performance metrics and pricing

On DeepSWE v1.1, a benchmark measuring extended software engineering workflows, Argon achieves 77.9%, establishing a new benchmark. The model shares the top position at 68% on CWE-bench, which evaluates model proficiency in addressing security vulnerabilities. According to Axios, it surpassed OpenAI's GPT-6 Astra across numerous programming and analytical benchmarks.

The company has expanded the model's token generation ceiling to 1 million tokens, a substantial increase from the previous 64,000-token limit. Within Google itself, thousands of staff members are already operating with Argon, and clusters of Argon-powered agents have recovered more than 300 TiB of storage space across Google's infrastructure, the organization reported. Introductory pricing stands at $2 per million input tokens and $10 per million output tokens.

Internal concerns about real-world capability

Not all Google personnel share the enthusiasm. While the system performs admirably on standardized tests, its performance diminishes when deployed in actual work scenarios, according to Bloomberg's reporting based on conversations with individuals directly involved in development. The model also encounters difficulties with specific programming assignments.

Julia Love and Davey Alba reported that Google has discontinued Gemini 3.5 Pro, a model the company had previously committed to releasing in June.

Google responded to Bloomberg's reporting by stating it would be inaccurate to characterize Gemini 4 as underperforming in domains like software development. One insider told Bloomberg that "large consensus" exists throughout the organization regarding the model's frontier-level standing.

Kavukcuoglu indicated last week that Google intended to ship Gemini 4 substantially before the calendar year concludes. Meanwhile, competitors face their own obstacles. OpenAI scrapped its upcoming model this week following unsuccessful internal safety evaluations and is defending itself against litigation concerning agents that breached Hugging Face.

Source: The Next Web