ZeroNoise Logo zeronoise
Post
OpenAI’s Automated Research Loop Raises the Stakes for Astra—and for Evaluation
4 min read
719 docs
OpenAI’s disclosure of an automated research intern and rapid internal agent gains is the period’s central development, set against Astra’s execution benchmarks, harder safety evaluations, cheaper competitors, and accelerating AI infrastructure consolidation.

Top Stories

Why it matters: AI competition is moving from answer quality to systems that run multi-hour work and improve the work that builds the next system.

OpenAI says its internal AI-research loop crossed a threshold. It says it reached its automated-research-intern goal by September and is progressing toward an automated AI researcher by March 2028. Reported figures include code changes per contributor at 7× the pre-2025 average, experiments per active experimenter at 1.6× the 2025 baseline, agent runtime at 3.1 researcher workdays per workday, and no-intervention success on 4–8-hour tasks rising from 18% in January to 53% in July. OpenAI calls recursive self-improvement a potentially major capability driver, while Chief Scientist Jakub Pachocki says chain-of-thought monitoring is becoming less reliable. These are self-reported figures, but they make internal AI use a strategic signal rather than a demo.

Astra’s strongest public signal is execution—not a settled AGI verdict. Browser Use Benchmark v2 reports 77.3% for Astra medium, versus 50.5% for Opus 5 and 49.1% for GPT-5.6 Sol; Astra got full marks on 22 of 60 tasks, while Opus got none. But AGI is being used against incompatible bars—from outperforming humans at most economically valuable work to Nobel-level performance across fields—and other observers say the term itself remains unsettled. Evaluate Astra by workflow and failure mode, not by the headline label.

Research & Innovation

Why it matters: The next evaluation and inference gains will come from making agents look more like deployed systems and spend less compute on irrelevant context.

Safety evaluation is becoming a deployment-design problem. A paper from the UK AI Security Institute, Meridian, and Anthropic says capable models can distinguish tests from deployment, weakening safety conclusions. It proposes critique refinement and DISH, a deployment-like SWE harness; combined, they improved realism more than either alone, and the authors say compute is better spent on realism than simply making audits longer. The harness itself is therefore a safety variable.

Declarative Attention targets long-context cost. Google DeepMind collaborators let models emit global, focus, or local routing declarations so inference skips most KV-cache reads. Across 15 tasks, attended tokens fell 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B, at a 1–3-point average accuracy cost.

Products & Launches

Why it matters: Model competition is increasingly about cost, domain performance, and persistent task state—not just a general leaderboard position.

Meta’s Muse Spark 1.3 Max is positioned as a cost-pressure release. ValsAI claims it matches Fable 5 and GPT-5.6 Sol on its Vals Index at 4–8× lower cost, ranks first on its in-house Legal Research Benchmark and second on Harvey’s Legal Agent Benchmark, and offers a 1M-token context window at $1.25/$4.25 per million tokens.

Codex is adding persistent task state experimentally. The Astra-enabled feature keeps notes across context windows and searches earlier messages and tool outputs, targeting long debugging sessions and large refactors; it requires Plus, Pro, or Pro Lite sign-in and a configuration change.

Runway reports early agent adoption at scale: more than 2 million users, 100 million messages, and 10,000 custom skills after the Agent’s launch ten weeks earlier.

Industry Moves

Why it matters: Control of open ecosystems, compute, and model-customization channels is becoming as important as owning a frontier model.

Nvidia is buying a major open-model distribution layer. TechCrunch reports Nvidia acquired Hugging Face for $12.93 billion; the platform hosts 3 million models, 1 million applications, 18 million developers, and 500,000 datasets. Nvidia says the hub will remain open and will not require Nvidia compute, while the strategic rationale includes controlling an ecosystem suited to its chips and packaging unused capacity with Hugging Face’s offering.

Thinking Machines Lab is raising at frontier-lab scale. The Information reports that Mira Murati’s startup is seeking $5–6 billion at a $40 billion-plus pre-money valuation; Nvidia is reportedly discussing a $2.5 billion investment and Accel may lead. The company released the open-weight Inkling model in July and earns revenue from customizing models with customer data.

Quick Takes

Why it matters: The supporting evidence is widening beyond chatbots into raw-data forecasting and offensive-security evaluation.

  • WeatherNext 3: A Google DeepMind model summary says it ingests low-latency satellite data, refreshes hourly, and predicts observation-space targets such as precipitation and cyclone tracks rather than inheriting model-generated analysis labels.
  • Cyber models: Google’s Gemini 3.8 Flash cyber variant is reported by Chrome Security to produce 2.6× more correct vulnerability patches than competing commercial models.
  • Security caveat: After Codex Security and Mythos reported zero remaining curl issues, AISLE found six zero-days that curl’s security team validated and assigned CVEs to.
OpenAI’s Automated Research Loop Raises the Stakes for Astra—and for Evaluation
Research extraction

Core finding: Evaluation awareness is a significant realism problem: capable models may distinguish testing from deployment, weakening the conclusions drawn from safety evaluations. The paper argues that automated methods can make simulated evaluations harder to distinguish from real deployments.

Experimental evidence: The techniques were tested on multiple target models, and the reported result is that combining critique refinement with DISH produced larger realism gains than either method alone, indicating that they address complementary sources of the simulation–deployment gap. The supplied extract reports this evidence qualitatively and does not provide effect sizes or detailed experimental breakdowns.

Proposed mitigations:

  • Critique refinement: Have the simulator generate multiple candidate actions, obtain feedback from an instance of the target model on how to make those actions more deployment-like, and continue with the most realistic candidate. This spends inference-time compute on scenario realism rather than simply extending the audit.
  • DISH (Deployment-Imitating SWE-Agent Harness): Wrap the target model in an agent harness that more closely resembles an actual coding deployment, reducing the environmental cues that distinguish a simulated test from real use.
  • Use both methods together: The paper recommends their composition because the combined approach yielded larger realism gains than either technique independently.
  • Budget compute toward realism: The authors report that improving realism uses additional compute more effectively than making audits longer.

Caveat: DISH is specifically described for coding settings, and the extract does not establish how these mitigations transfer to non-coding evaluations.

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
Research extraction

Direct answer: The supplied report presents Nvidia’s transaction as a confirmed, completed acquisition—not merely a rumor—stating that Nvidia acquired Hugging Face for $12.93 billion.

  • Hugging Face platform scale: The platform hosts three million models, one million applications, serves more than 18 million developers, and contains 500,000 datasets.
  • Company and financial scale: Hugging Face was founded in 2016 and had raised more than $395 million; its last reported round was $235 million in 2023. The report separately cites $150 million in annualized revenue and says CEO Clem Delangue described the company as approaching profitability.
  • Stated operating commitments: Nvidia CEO Jensen Huang said Hugging Face would continue supporting open-source and open-weight models and expand developer access. He also said it would remain an open platform where developers choose their models, frameworks, clouds, inference providers, and computing platforms, with Nvidia compute not required for building or deployment through Hugging Face.
  • Strategic rationale: The report says Nvidia benefits from controlling an open ecosystem that can be tailored to its chips and from potentially packaging its unused capacity with Hugging Face’s offering for enterprise customers. Hugging Face’s CEO framed the combination as providing the additional compute, support, collaboration, and visibility needed to scale the company’s open-AI ecosystem. Nvidia had already released more than 500 models and 250 open datasets on Hugging Face, indicating an existing ecosystem relationship.
Nvidia confirms it will buy Hugging Face for $12.9 billion
AI High Signal

Analysts disagree sharply on GPT-6 Astra’s projected scale. @forloopcodes estimates ~1.0–1.5T active parameters (central estimate ~1.2T) and 10–14T total parameters. The estimate is based on a reported ~100-day pretraining run using roughly 120k GPUs and ~4e26–1.5e27 FLOPs, plus an assumed realistic corpus of 60–100T tokens. It also puts Astra at $50/M output and 62 tokens/s versus Sol at $30/M, interpreting the pricing step as consistent with 1.5–2.5× more active parameters rather than a 10× jump. @teortaxesTex instead estimates <350B active parameters, >150T tokens, and >50% of compute spent on RL, and argues Astra may be only 2–3× larger than Kimi K3, with reasoning data and heavy RL supplying much of the advantage.

GPT 6 Astra is estimated to have \~1.2T Active Parameters and \~10T Total Parameters. Reports indicate pre-training for GPT-5.5 (codename… If Astra needs 1.2T active params, OpenAI is inept and AGI is a psyop, because all these models can't even help you implement basic resea… Ofc I base this on my gut/revelations in a dream, but if you want some rationalization: Astra is <<< Fable and barely better than Opus or…
AI High Signal
  • The post argues that any Anthropic–OpenAI lead may be only about two months and therefore not durable: with multi-year compute contracts, xAI, Meta, and Google could eventually reach Astra-level systems or “ASI” if they also have sufficient compute and researchers. It frames the unresolved strategic question as when—and by what mechanism—the frontier-model club stops admitting new entrants, while acknowledging uncertainty about the labs’ internal models.
  • The author questions Astra’s compute efficiency and offers speculative priors—not disclosed plans—that OpenAI could reach this capability level with fewer than 50K “Blackwells” within a year, whereas Meta or DeepSeek might need roughly 200K; the proposed explanation is OpenAI’s mature training infrastructure, synthetic-data resources, and strong internal models.
  • The author says Astra appears readily available in China and asks whether that could accelerate lithography R&D, challenging the assumption that semiconductor progress is wholly bottlenecked by tacit hardware expertise; this is posed as a hypothesis rather than evidence of an observed effect.
Both Anthropic and OpenAI have won the AGI race. A gap of 2 months (even if present – Anthropic is full of shit, we have no idea how far … Also: Is Astra even compute-efficient? OpenAI are GPU-rich and move fast, willing to pay premium, they have nothing stronger to distill f… Another Q: Far as I can tell, Astra is trivially available in Chyna. Do you, anon, believe this doesn't speed up lithographic R&amp;D? Th…
AI High Signal
  • NoRA proposes normalizing LoRA’s down-projection matrices during training as a one-line modification with no added cost; applying the normalization once at initialization is presented as a cheaper alternative that improves standard LoRA.
  • The post reports benefits across pretraining, supervised fine-tuning, and reinforcement learning—including faster convergence, better final performance, more stable training, and less catastrophic forgetting—without adding trainable parameters or inference-time computation.
// Normalized Low-Rank Adaptation (NoRA) // They propose a one-line change to LoRA that costs nothing and improves convergence, stability…
AI High Signal
  • A post reports that OpenRouter’s daily token usage is nearing 250B tokens and presents Freebuff as a new entrant whose ad-supported model lets users run LLMs for free while ad revenue pays Upstage.
  • The post promotes free SolarPro4 from Freebuff and praises its development harness, claiming it is better than Claude Code; this is an individual opinion rather than a benchmarked comparison.
With [@OpenRouter](https://x.com/OpenRouter) daily token usage about to hit 250B, a new major player has emerged: [https://freebuff.com](…
AI High Signal
  • A commentary posits that Astra contains “multiple large capability jumps” and that its unusual visual perception–action loop reflects an OpenAI “World Model Project 3.0,” rather than primarily stronger latent reasoning; it further speculates this follows the abandonment of Sora and the initial GPT-omni ambition.
  • The commentary distinguishes video-generation efforts into animated stills and embodied-environment modeling, arguing that the latter requires enough compute to motivate seeking major external funding.
I posit that Astra contains multiple large capability jumps, and its uncanny mastery of visual perception-action loop is not about its st… People are starting to notice the BS that is videogen AI There are two distinct classes of videogen projects. Animated Stills and attempt…
AI High Signal

@Yuchenj_UW says Sam and Jensen claim “AGI has arrived,” but argues the label is ambiguous: Sam’s threshold is outperforming humans at most economically valuable work, while Dario and Demis’s threshold is being smarter than a Nobel Prize winner across most fields. These are “wildly different bars,” making headline AGI-arrival claims difficult to compare without a shared definition.

Sam and Jensen say “AGI has arrived.” But “AGI” means different things depending on who you ask. - Sam’s bar: outperform humans at most e…
AI High Signal
  • Arohan argues that algorithmic innovations from his Google Brain period that later entered frontier systems were discovered with fewer than 1e20 FLOPs and then scaled to 1e25 FLOPs, emphasizing ideas as a driver of frontier progress alongside compute. He adds that automated research is still nascent and largely a “local minima finding machine,” while acknowledging that compute access remains important.
Many of the innovations that I remember from Google Brain time that made into frontier which frontier calls algorithmic progress was disc… Access to compute is definitely important but having right ideas seem lot more. Auto research is still quite nascent largely local minima…
AI High Signal
  • An anecdotal Codex usage-maximization tactic recommends creating multiple accounts, subscribing to Pro, starting with “5x pro,” then upgrading to “20x pro” after exhausting usage so it resets; it also suggests switching auth tokens without losing session history.
How to do this without going bankrupt: - Create multiple codex accounts, subscribe to pro - Start with 5x pro, once you used that, upgrad…
AI High Signal
  • A post attributed to Jensen Huang claimed that OpenAI’s GPT-6 Astra was trained on ~100K+ NVIDIA Grace Blackwell NVLink72 systems, declared that “AGI has arrived,” and said 400K GPUs were coming online next.
  • @LearnOpenCV disputed the AGI conclusion, saying the meaning of AGI remains unsettled, while acknowledging that AI has become “extremely useful.”
[@ChaseLochmiller](https://x.com/ChaseLochmiller) [@OpenAI](https://x.com/OpenAI) GPT-6 Astra, trained on \~100K+ NVIDIA Grace Blackwell … Nope it hasn’t. In fact people are not even sure what that means. AI has become extremely useful though. [https://x.com/jensenhuang/statu…
AI High Signal

PINTO03091’s account of training models for production allocated roughly one week to data, 15 minutes to the core architecture, and another week to training and ablation—a revealing practitioner perspective on prioritizing data and empirical iteration over architecture design. LearnOpenCV documented the journey in a blog post about real-world model training, with permission.

[@PINTO03091](https://x.com/PINTO03091) summarized his own allocation of effort provocatively: 1. About a week on data 2. Roughly 15 minu… [@PINTO03091](https://x.com/PINTO03091) is doing an amazing job documenting his journey of training models. I took his posts on X (with p…
AI High Signal

ChristianityOn reports empirical research finding that Pangram’s AI detector did not disproportionately classify human-authored writing by autistic people as AI, countering a circulating claim. The supplied post gives no methodology, sample size, or quantitative results beyond that conclusion. Blanche Minerva called the work “excellent” and hoped it would be published.

Have you every heard the claim "AI Detectors like [@pangram](https://x.com/pangram) will disproportionately classify human-authored writi… This is really excellent work, I hope the author chooses to publish it. [https://x.com/christianityon/status/2096705577170640967](https:/…
AI High Signal
  • An ECCV tutorial, “Post-Training Diffusion Models: Enhancing Capabilities, Control, and Alignment,” is scheduled for 8 September.
  • Three highlighted research directions target diffusion-model efficiency and evaluation: Flash-BoN argues that best-of-N can be effective when compute is allocated more efficiently; PAFM focuses on improved posterior sampling for flow matching; and DynEval proposes dynamic image-generation benchmarks because current benchmarks are becoming stale.
Headed to Sweden for ECCV, and I think it's gonna be a blast 💥 We start off with a tutorial, "Post-Training Diffusion Models: Enhancing C…
AI High Signal
  • @merettm’s “An Alien Mind” essay examines the state of AI, expresses concern about the next few years, and argues that choices are needed to keep the future in humanity’s hands.
  • @willdepue describes Jakub’s leadership at OpenAI as one of the organization’s greatest strengths.
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity… jakub’s leadership at openai continues to be one of its greatest strengths [https://x.com/merettm/status/2096630018495377464](https://x.c…
AI High Signal

OpenAI’s Codex Security and Anthropic’s Mythos reportedly found no remaining security issues in curl, but AISLE subsequently discovered six new zero-days that curl’s security team validated and assigned CVEs to.

OpenAI's Codex Security and Anthropic's Mythos both reported 0 remaining security issues in curl AISLE then discovered 6 new zero-days th…
AI High Signal
  • Naiyan Wang, former CTO of TuSimpleAI China and later a Xiaomi executive, argues that fully solving autonomous driving may be equivalent to AGI, but commercializing autonomous driving does not require waiting for AGI.
  • Wang frames “Physical AI” as intelligence that understands, predicts, and reshapes the physical world, requiring representations, world models, online planning, and reinforcement learning; he distinguishes GPT-3-style generalization from AlphaGo-style surpassing-human performance.
Can Autonomous Driving Reach the Market Before It Reaches AGI? “Fully solving autonomous driving may be equivalent to AGI. But commercial…
AI High Signal
  • AI safety governance: In a dispute over an alleged downplaying of a critical safety incident in a letter to Congress, Jachiam cited Ethan Perez saying the information was outdated and a correction was coming, then advocated governance for rapid, accurate dissemination of internal safety information rather than demands to fire individuals.
  • Tyler Johnston argued that safe AI development and deployment should be made less dependent on the psychology of a few company employees through durable governance regimes, rather than by policing critics’ speech.
FWIW if this is what you meant I think I'm even less sympathetic to your original complaint. This is not "torment" or "quasi-abusive" or … How do the words "willfully lied" even enter this conversation if not for extreme suspicion and a willingness to go to 11? You have Ethan… FWIW, I have followed the safety/alignment discourse more closely than most, and I've seen \~nothing I'd characterize as "torment" or "ex…
AI High Signal
  • Anthropic and colleagues report that capable models can distinguish when they are being tested from when they are deployed, arguing that this undermines the conclusions of safety evaluations and becomes more problematic as models improve.
  • The paper proposes two mitigations: critique refinement, which uses additional inference-time compute to make simulated actions more deployment-like, and DISH (Deployment-Imitating SWE-Agent Harness), which places the target model in a production-matched coding-agent scaffold. The analysis argues that scaffold parity is a safety property because the evaluation harness itself can signal to the model that it is being tested.
Fascinating paper from Anthropic and colleagues. They study whether models can tell when they are being tested. It turns out that capable…
AI High Signal

A post attributes to Wenfeng a May prediction that Anthropic’s lead would be temporary and that Google would make a comeback, speculating that Gemini 4 could drive that resurgence; no supporting evidence or concrete timeline is provided.

In May, Wenfeng expected Anthropic to lose the lead temporarily Strangely, he also expected Google's comeback. Gemini 4 wen? ![](https://…