ZeroNoise Logo zeronoise
Post
Google Puts Gemini 4 Argon Back at the Frontier, With a Gated Launch to Cyber Defenders
•
5 min read
• 1021 docs
Gemini 4 Argon matches GPT-6 Astra on independent indexes at launch-discount prices, though access is limited for now. OpenAI attributes a reasoning-distillation campaign to people linked to Moonshot, and agentic and open-model releases continued.

Gemini 4 Argon: Google rejoins the top tier, but on a short leash

Google DeepMind's Gemini 4 Argon is its first proprietary model above the Flash class in more than seven months. Independent benchmarks put it level with the best models available. On the Artificial Analysis Intelligence Index, Argon at high reasoning scores 53. That matches GPT-6 Astra (max), is one point ahead of GPT-6.1 Sol (max), and is 23 points above Gemini 3.1 Pro Preview . Other results:

  • Google's own comparisons: first on 13 of 19 published benchmarks against Astra and Opus 5.5. That includes DeepSWE (77.9% vs. 74.2% for Opus 5.5) and the Vals Index (68.9% vs. 67.0%) .
  • Arena: #1 in Text Arena with 1,525 points and #8 in WebDev .
  • Agentic and accuracy tests (Artificial Analysis): #1 on AutomationBench-AA at 77.5%, but behind Sonnet 5.5, Opus 5.5 and Astra on Terminal Bench 4 at 57%. Its hallucination rate is 15%, compared with 51% for Astra. Its raw accuracy is 50%, 13 points below Astra .

Price and access. Standard pricing is $4/$20 per million input/output tokens. A launch promotion halves that to $2/$10, and cached input gets a 95% discount. Argon has a 1M-token context window, and a new Long Decode Continuation feature lets reasoning run up to 1M output tokens . At the promotional price, Artificial Analysis puts it at $1.99 per Intelligence Index task, 60% of Astra's cost but 2.7× GPT-6.1 Sol's. The cost would rise to $3.98 once the discount ends, and Google has not said when that will be. The savings come from cheaper tokens, not from using fewer of them: Argon averages 62k output tokens per task against Astra's 27k . The model is not publicly available yet . Sundar Pichai says it is with the US government and with trusted cyber defenders in the Fairwind Program, and that wider availability will come "as soon as we can and as safely as we can" .

Internal use. Google says Argon agents freed more than 300 TiB of memory across its data centers. They are also migrating more than 800,000 lines of C/C++ kernel code to Rust . Vahab Mirrokni credits the model with helping complete the "CK conjecture" and says further results are queued for release .

Skepticism. One commenter flagged Argon's 19.6% on Harvey's Legal Agent Benchmark, below Muse Spark 1.2's listed 25.42% . @teortaxesTex expects Argon to disappoint in practice. He still thinks it matters strategically because it is "a completely new pretrain" that Google can iterate on .

The release fits a broader trend. Artificial Analysis says the gaps between successive record scores on its index have shrunk from 239 to 126 to 99 days .

OpenAI attributes a reasoning-extraction campaign to people linked to Moonshot

OpenAI published a report on a coordinated campaign to extract its models' hidden reasoning. It attributes "a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi." It adds that it is unclear whether all the operators were one actor . According to a summary of the report, OpenAI logged 16,000 extraction attempts from more than 4,000 users in two days, and found related activity across more than 15,000 users. Operators moved encrypted reasoning between conversations and prompted the model to reveal it .

Separately, academic researchers report that two months after they disclosed their own extraction attack, they could still pull reasoning from Astra and Sol 6.1 through third-party API providers. They have told OpenAI and Anthropic . One author said patches are hard to roll out across every product version and third-party vendor .

GPT-6.1 Sol: demand and cost data

Codex lead Thibault Sottiaux says GPT-6.1 Sol is OpenAI's "most demanded model pretty much ever" across both the API and subscriptions. After heavy load, OpenAI has added capacity and expects serving speed to nearly double . ARC Prize verified scores of 94.2% on ARC-AGI-2 and 96.4% on ARC-AGI-3 with the provider's adapter harness. On ARC-AGI-3 with the standard harness, it scored 52.7% . Artificial Analysis says it costs about 30% less per task than GPT-6 Sol. The savings come from using fewer turns and a cheaper cache-read price, not from a price cut .

Biosecurity and policy

  • SynthID Bio. Google DeepMind synthesized AI-designed proteins that are both functional and watermarked, published the work in Nature, and is open-sourcing the tools .
  • FTC (unconfirmed). An X account claims the FTC has confirmed an investigation into OpenAI, Anthropic and METR, and is drafting civil investigative demands about the dangers those labs have warned of . No primary source was reviewed.
  • Brockman. The NYT reports that Greg Brockman is no longer making a second $25M donation he had promised to the super PAC Leading the Future .
  • SI czar (rumor). Rumors say DNI Jay Clayton will be named White House superintelligence czar. He told CNBC that superintelligence "is a national security issue" .

Models, agents and infrastructure

  • Runway Praxis-1. An open-weight "World Action Model" that uses video pretraining to produce robot policies for any embodiment. Runway is testing it with Noble Machines, Standard Bots and Ultra, and plans to release the weights in the coming months .
  • Ling-3.1-flash. About 560B total parameters, about 25B active, with up to 1M-token context. Open-sourcing is planned .
  • Superhuman Stratego. A Nature paper reports the first superhuman Stratego AI, built with general reinforcement-learning and test-time-compute methods for games with hidden information .
  • DeepSeek. Released an open-source kernel toolkit for Huawei Ascend, with TileLang support optimized for the Ascend 950. It is framed as a CUDA alternative . It also shipped a Harness v0.2 desktop preview for macOS and Windows. DeepSeek says Harness is the most-used coding agent among users of its official API .
  • Embeddings. Perplexity's open pplx-embed-v2-context-9b-preview claims state-of-the-art results on ConTEB and turbopuffer's context-bench . Cohere's Embed 5 Pro leads the models Cohere itself measured on ViDoRe V3 .
  • Hardware. Cognition is the first customer running NVIDIA Vera Rubin, on CoreWeave. It reports about 4.8× the token throughput of GB200 at the same decode speed . Synopsys signed a multi-year deal worth more than $1B to supply silicon IP for Amazon's custom chips .
  • Video. On Artificial Analysis' new text-to-video benchmark, open-weights MiniMax H3 ranks #3, statistically tied with Seedance 2.5 at about one-seventh the price .
  • Factory vs. Cognition. Factory removed Chris Degnan as board observer and advisor. It says he had been holding formal, recurring talks with Cognition executives while advising Factory on confidential matters. It also accuses Cognition engineers of faking interviews to extract product information .
  • Gemini app. Skills will replace Gems, starting in November for personal accounts .
Google Puts Gemini 4 Argon Back at the Frontier, With a Gated Launch to Cyber Defenders
AI High Signal
  • Google’s Mirrokni announced that Gemini 4 Argon is live. He said Google had used it with internal agents for large software-engineering, math, and RSI tasks, and that the model helped complete the CK conjecture; other results were still awaiting release.
  • Artificial Analysis scored high-reasoning Argon 53 on its Intelligence Index, tying GPT-6 Astra (max) and exceeding GPT-6.1 Sol (max, 52); it also ranked first on AutomationBench-AA at 77.5%. Its reported hallucination rate was 15%—lowest among models scoring 45+—but its accuracy was 50%, versus GPT-6 Astra’s 63%.
  • Argon has a 1M-token context window and accepts text, image, video, and speech input; Long Decode Continuation can extend reasoning to 1M output tokens without request timeouts.
  • Artificial Analysis put the cost at $1.99 per Intelligence Index task under the current 50% promotion (60% of GPT-6 Astra’s cost), rising to $3.98 at standard pricing; Argon averaged 62k output tokens per task versus Astra’s 27k, and Google had not announced when the promotion would end. Mirrokni called the model live, while the benchmark post described a selected-user rollout and said it was not publicly available.
Gemini 4 Argon is finally live! Thanks to the amazing Gemini team. Proud of our team collaborating on many parts of Gemini. I'd like to p… Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted …
AI High Signal
  • Mirrokni announced that Gemini 4 Argon is live; the accompanying Artificial Analysis report described it as rolling out to selected users and not yet publicly available. Artificial Analysis scored its high-reasoning version 53 on the Intelligence Index, matching GPT-6 Astra and exceeding GPT-6.1 Sol by one point; that is 23 points above Gemini 3.1 Pro Preview’s score of 30.
  • Artificial Analysis reported strong agentic results, including a first-place 77.5% on AutomationBench-AA, and a 15% hallucination rate—the lowest among models scoring 45+ on its index. It also noted a trade-off: Argon’s 50% accuracy was below GPT-6 Astra’s 63%.
  • At the reported 50% launch discount, Argon costs $1.99 per Intelligence Index task versus $3.26 for GPT-6 Astra; at standard pricing, the task cost rises to $3.98, and the promotion’s end date was unconfirmed. Argon has a 1M-token context window, while Long Decode Continuation allows reasoning to run up to 1M output tokens without request timeouts.
  • Mirrokni said internal agents use the model for large software-engineering, math, and RSI tasks, and credited it with helping complete the CK conjecture and other results that were still awaiting release.
Gemini 4 Argon is finally live! Thanks to the amazing Gemini team. Proud of our team collaborating on many parts of Gemini. I'd like to p… Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted …
AI High Signal

The Human Knowledge Compression Contest (Hutter Prize) reported three winning submissions within one month, nearly 10% total improvement and about one-third the zip-file size; the post called this the largest progress in the contest’s 20-year history. Vladimir “astOwOlfo” Ivanov was identified as the contest’s eleventh winner.

The Human Knowledge Compression Contest (widely known as the Hutter Prize) is alive and kicking. This year has seen the largest progress … Congrats to Vladimir "astOwOlfo" Ivanov the eleventh Winner of the Hutter Prize ![](https://pbs.twimg.com/media/HThfSViWgAA3GWi.png) [htt…
AI High Signal

A discussion argues that much of what is currently called “RSI” is better described as “Recursive Technological Improvement”: improvement is not confined to one entity improving itself, and techniques diffuse rapidly. The follow-up contrasts FOOM’s imagined program recursively editing itself with what the author characterizes as today’s paradigm of spawning and breeding populations.

speaking of Comprehensive AI Services, most of what people currently call "RSI" is better described as Recursive Technological Improvemen… The issue with RSI is with the notion of «self». FOOM assumed a mind substrate close to Elisp: virtually no function approximation, just …
AI High Signal

@reach_vb dates the Codex app’s release to February 2, 2026, and says it changed how millions interact with agents and inspired a generation of apps .

2nd February 2026! An app that changed the way millions of people interact with agents and inspired a generation of apps ![](https://pbs.… Your reminder that the codex app came out in February
AI High Signal

DeepSeek is reported to offer a $200 Pro coding plan with no five-hour, weekly, or rate limits; usage stops when the $200 is spent, and the post says other budget amounts are available.

DeepSeek does offer a $200 Pro coding plan, like OpenAI no 5-hour limits, weekly limits, rate limits, or really any limits, you just run …
AI High Signal

KV-streams preserves the KV cache during agentic compaction rather than flushing it, avoiding re-prefill costs for both the inference engine and trainer; the authors say it works with any compaction method. In SWE tasks, they report matching the performance of re-prefill compaction and full context in half the time, after first experimenting with text-based games.

Want to train your SWE agents 2× faster 🏎️💨? We introduce KV-streams, which, rather than flushing the KV cache on agentic compaction, we p…
AI High Signal

The post reports training a new type of model for robots and winning another hackathon this year, claiming a 100% win rate so far . It says more details are forthcoming .

distilled astra, trained a new type of model for robots and won yet another hackathon this year (100% win rate so far woohoo) will share …
AI High Signal

jevgrep is a research-agent CLI powered by jev from @typesafeai, presented as reducing coding-agent costs by 40% (verified on SWE-bench). Its built-in skill guides coding agents to use jg for context collection.

Introducing jevgrep - a research agent CLI powered by jev from [@typesafeai](https://x.com/typesafeai) that reduces your coding agent cos…
AI High Signal

Attendees at the Runway AI Summit in San Francisco received an early preview of Project Continuum; more information was promised, but the post provides no details about the project.

Wrapping up the Runway AI Summit in San Francisco, many people asked me what I am most excited for. Attendees of the event got a special …
AI High Signal

A computer-use-agent speedrun reports a 4.4× speed gap between Astra and Kimi K3 despite the two achieving the same score.

cua-speedrun Speedrun for computer use agents. There is a massive gap in speed e.g. 4.4x between Astra and Kimi K3 despite achieving the …
AI High Signal

The Looped Diffusion Transformer repeatedly runs shared Transformer blocks within each denoising step and is reported to surpass a model 6.5× larger across text-to-image benchmarks while using 4.9× less inference compute.

Looped Diffusion Transformer - Scales computation by repeatedly running shared Transformer blocks within each denoising step - Surpasses …
AI High Signal

Google DeepMind introduced Gemini 4 Argon, a new frontier model for complex coding, enterprise knowledge-work, and cybersecurity-defense workflows; it is rolling out to a set of trusted testers through the Fairwind Program.

Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cyber…
AI High Signal

Hermes now supports adding full new languages across all its surfaces via Plugin, beyond the 16 languages already supported.

You can now add full new languages across all of Hermes' surfaces via Plugin in addition to the 16 already supported
AI High Signal

Airbnb’s updates include AI search using users’ own words, described as a first step toward an agentic-AI interface the company says is coming soon.

Hey everyone, we made some updates based on feedback from you. We’ll keep shipping, so I’ll post more updates soon - connect with your fr…
AI High Signal

Recent score milestones on the Artificial Analysis Intelligence Index arrived at progressively shorter intervals—239, 126, then 99 days; the current leader scores 57.6, signaling a faster pace of progress on this index.

The pace of AI progress is picking up. The time between recent score milestones on the Artificial Analysis Intelligence Index decreased f…
AI High Signal

Nous used NVIDIA NeMo Relay to collect execution traces for Hermes Agent; a walkthrough runs two scenarios, shows agent calls and retries with full traces in Arize Phoenix, and describes evaluating fixes using traces and task results across repeated runs. Teknium said the additional data helped optimize the agent harness.

Your agent got the right answer. How much work did it take to get there? We worked with [@NousResearch](https://x.com/NousResearch)/@Tekn… The more data you have to look at the better you can use it to optimize a harness, thanks to nemo relay we had a lot more to work with! […
AI High Signal
  • Sakana AI cofounder and CEO David Ha argues that as frontier scaling slows, AI’s next phase may depend on orchestrating multiple models according to their strengths rather than pursuing one dominant, ever-larger model; he notes that open-model gaps have narrowed to months and frontier inference costs can exceed the hourly cost of the people those models are meant to assist.
  • Ha frames sovereign AI as supply-chain resilience, not isolation: countries should build domestic capability to develop, tune, and operate AI while combining global resources and retaining alternative access if a provider cuts service.
Sakana AI共同創業者・CEOのDavid Ha(@hardmaru)の寄稿が [@NikkeiAsia](https://x.com/NikkeiAsia) に掲載されました。タイトルは「AIの未来はオーケストレーターにある」。 フロンティア企業が巨額の計算資源を投… I just published an op-ed in [@NikkeiAsia](https://x.com/NikkeiAsia) on why the future of AI belongs to the orchestrators. A school of fi…
AI High Signal
  • Artificial Analysis reports that GPT-6.1 Sol sets a new cost-efficiency frontier among OpenAI models at every effort level; at max effort it costs $0.72 per Intelligence Index task, versus $3.26 for GPT-6 Astra, $1.04 for GPT-6 Sol, and $1.99 for GPT-5.6 Sol.
  • OpenAI fixed an image-encoding bug in GPT-6 Luna and Sol. GPT-6 Luna (max) gained one Intelligence Index point, with reported gains on visually oriented evaluations including GDP.pdf (+2.4 points), MMMU-Pro (+4.1 points), GDPval-AA v2.1 (+71 Elo), and AA-Briefcase (+36 Elo); GPT-6 Sol showed negligible improvement.
GPT-6.1 Sol pushes the cost efficiency frontier for OpenAI models. OpenAI has also fixed an issue reading multimodal input which improves…
AI High Signal

RLance Martin said he will update a Claude skill that builds a viewer for eval examples to first walk users through the data, which he agreed could help them prioritize which evals to write.

thanks [@HamelHusain](https://x.com/HamelHusain) + [@isaac_flath](https://x.com/isaac_flath)! the latest skill does instruct Claude to bu…