ZeroNoise Logo zeronoise
Post
AI Progress Is Now Bounded by Cyber Risk and Harness Reliability
17 hours ago
3 min read
498 docs
This brief covers a reported Astra safety pause, new evidence of harness-dependent agent behavior, and workflow-specific products and strategic deployments.

Top Stories

Why it matters: Safety gates and integration reliability are becoming as consequential as raw model scores.

OpenAI reportedly pauses Astra scaling. @dl_weekly reports that OpenAI paused its largest frontier RL run and slowed scaling after preliminary evidence that upcoming Astra may cross the “Critical” cybersecurity threshold; it is hardening research-environment security, monitoring, and alignment before proceeding. If borne out, the operational signal is that capability-risk evidence is changing the training schedule itself.

Agent capability is not portable by default. An analysis shared by @ZhihuFrontier says the same DeepSeek V4 Pro weights can approach their ceiling under DSH’s minimal preset and degrade in standard or third-party frameworks. It identifies format, context-structure, and control-flow overfitting; production failures included unfamiliar tool names pushing shell workarounds, truncated results causing abandonment, and extra MCP tools weakening stopping. Randomized harnesses narrowed worst-case variance, so teams should track the worst integration—not just the mean—and version the harness as a model dependency.

Research & Innovation

Why it matters: The highest-leverage technical work is increasingly instrumenting the agent loop around the model, not only changing model weights.

Agent Lightning v1.0 from Microsoft uses a roughly 3,500-line endpoint proxy to connect an existing harness to RL without rewriting the agent. With 6K training examples, Qwen3.5-9B moved from 41.8% to 56.4% on SWE-bench Verified; the paper warns that retokenization, sample merging, advantage calculation, loss normalization, and scheduling can silently corrupt gradients at the harness boundary.

Netflix is operationalizing LLM judges. Its system evaluates hundreds of thousands of recommendation explanations weekly for millions of mobile members; the judge has birth, training, deployment, and continuous-monitoring phases. A five-week A/B test over tens of millions shifted viewing toward previously unwatched content and increased browse-to-play sessions without quality-related takedowns.

Structured conflict beats agent proliferation. Adversarial Review uses a coder, reviewer, and critic; it beat a five-agent baseline on LiveCodeBench. On SWE-PRBench, explicit disagreement addressed a failure mode where agents converged without enough evidence and produced the highest F1 among tested methods.

Products & Launches

Why it matters: Capability is being packaged inside multimodal creation tools and domain-specific development workflows.

Wan 3.0 is live on fal with native 30-second, single-pass clips, omni-reference input across text, images, audio, video, web pages, and documents, and improved real-world motion. fal exposes text-to-video, image-to-video, and reference-to-video endpoints.

MongoDB Agent Skills and Plugins package structured guidance on schema design, indexing, query patterns, connection management, and AI retrieval for Claude Code, Cursor, Gemini CLI, and VS Code. MongoDB’s MCP server manages authentication and scopes what agents can access, with configurable controls such as disabling tools; the plugins join that connectivity layer to the skills.

Industry Moves

Why it matters: Strategic AI spending is pairing specialized infrastructure with mission-specific agents.

Sakana AI enters Japan’s intelligence workflow. Sakana’s official release says it signed a Ministry of Defense contract to research and demonstrate AI functions for integrated intelligence analysis. The project applies its AI-agent technology to more efficient collection, better analysis, and systematic information management for decision support; Sakana places it at the strategic-intelligence level, alongside earlier C2 research.

Etched’s financing is also a customer-validation signal. TheTuringPost reports $1 billion raised in 26 days and a valuation jump from $10.3 billion to $21 billion; Jane Street tested the hardware, installed the first rack, and led a $700 million round. Etched says its system is now architecture-agnostic across Llama, DeepSeek, Qwen, and Mamba, though the report treats that flexibility as an unresolved question rather than a settled result.

Quick Takes

Why it matters: Pricing and budgeted throughput are becoming visible product differentiators.

  • Together Compute says a $100 budget yields about 17 solved DeepSWE tasks with GLM-5.3 versus 3 with Fable 5, despite near-equal first-try performance; it is a vendor-reported benchmark, but it makes cost per completed task hard to ignore.
  • Codex’s usage reset has propagated, with fixes landed for the previously identified usage problems and more work promised.
  • A monitored update says DeepSeek now applies off-peak API rates all day on weekends.
AI Progress Is Now Bounded by Cyber Risk and Harness Reliability
Research extraction

Direct answer: MongoDB has launched official MongoDB Agent Skills and bundled plugins, with support for Claude Code, Cursor, Gemini CLI, and VS Code, and the bundle combines the MongoDB MCP Server with Agent Skills . Skills can also be installed for most coding agents via the Vercel Skills CLI or by manually cloning the GitHub repository , and the MCP Server plus skills are described as installable for Claude Code, Cursor, Gemini CLI, or other AI development tools .

Capabilities covered: the skills provide structured instructions, best practices, and resources across the development lifecycle, including schema design, performance optimization, and advanced capabilities like AI retrieval ; they encode expert guidance on schema design heuristics, indexing strategies, query patterns, and operational safeguards . The initial release spans the full application development lifecycle on MongoDB, from connection management and schema design to guidance on implementing advanced capabilities , and addresses common pitfalls such as over-normalizing schemas, underusing compound indexes, and misusing indexes and search indexes . The skills themselves follow the open Agent Skills standard introduced by Anthropic, which the source says is adopted by Claude Code, Cursor, Codex, and more .

Access-control and MCP features: the MongoDB MCP Server serves as the connectivity layer, manages authentication, and defines exactly what agents can access and do; combined with MongoDB's native authorization it ensures agents operate with only the permissions they need, and governance is provided through configurable controls such as disabling specific tools . While some skills work independently, others work in conjunction with the MCP Server for workflows that require it . MongoDB also encourages teams to build custom skills to codify internal conventions and workflows . The source notes that MCP Server and Agent Skills are designed to work best together, giving guardrails and control .

Introducing MongoDB Agent Skills and Plugins for Coding Agents
Research extraction

Direct answer: Sakana AI's announcement states it concluded a Ministry of Defense contract on July 29, Reiwa 8 for a 'survey/demonstration' (調査・実証) of AI functions needed for integrated analysis work — a research and demonstration effort, not deployment.

  1. Official scope and nature. The contracted work is '総合分析業務に必要なAI機能の調査・実証' (survey/demonstration of AI functions required for integrated analysis work). It targets the '総合分析業務' carried out at 情報本部 (the MOD's intelligence organization) and will demonstrate improvements in '情報収集の効率化', '分析能力の向上', and '体系的な情報管理' (information collection efficiency, analysis capability, and systematic information management). The project aims to advance 'social implementation' of decision-supporting AI functions, but the contract title and description use 調査・実証, indicating a research/demonstration effort.

  2. Intended AI-agent capabilities. The project will use 'Sakana AIの高度なAIエージェント技術' (advanced AI-agent technology) and applies cutting-edge AI to intelligence analysts' work, with the goal of fundamentally strengthening the integrated analysis work. The release does not disclose specific agent architecture, target tasks, or technical deliverables.

  3. Positioning versus earlier defense work. Sakana AI states it received a contract this March from the ATLA Defense Innovation Science and Technology Institute for foundational technology research toward C2 system enhancement. It contrasts that work (integrating unit/sensor information for rapid, accurate command) with the new contract, which targets strategic intelligence supporting overall security policy, from unit-operations command/control to information analysis supporting national-level policy decisions.

  4. Gaps and uncertainty. The only supplied source is Sakana AI's press release; no contract value, contract period, milestones, or third-party/MOD confirmation is included. 'Social implementation' is an aim, not evidence of a deployment contract.

Sakana AI、防衛省から「総合分析業務に必要なAI機能の調査・実証」を受注
AI High Signal

@obsxrver claims Amazon, one of the biggest robotics investors, deploys zero humanoid robots and invests <1% of its robotics spend in the technology . @teortaxesTex counters that China already leads in EVs, drone ships, flying drones, industrial robots, logistics robots, and humanoid robots, and argues mass-produced humanoids will win on cost via Wright's law — with 24/7 operation enabling amortized labor costs under $20/hour . He also argues humanoids will optimally perform many functions, though not warehouse pallet moving; @obsxrver disputes this, saying robotics form should be engineered to function .

[@teortaxesTex](https://x.com/teortaxesTex) He’s completely right. Let’s look at one of the biggest investors in robotics and industrial … [@obsxrver](https://x.com/obsxrver) This is retarded China can walk and chew bubblegum at once. They dominate in EVs, drone ships, flying… the fundamental disagreement I have is with \*you\*. you assume flexible manual labor is not worth a generalist robot body R&D. As you ca… [@obsxrver](https://x.com/obsxrver) They \*will\* optimally perform a whole host of functions, just not moving pallets in an Amazon wareh… [@teortaxesTex](https://x.com/teortaxesTex) I don’t disagree with the China claim but that’s kind of a non sequitur tho. I see humanoid r…
AI High Signal
  • teortaxesTex argues China now controls nearly the entire humanoid-robotics supply chain, forcing US builders to smuggle parts in suitcases; most relevant talent sits in dense Chinese clusters, and low-cost rapid prototyping makes China's iteration speed "probably impossible to match" .
  • Claims every promising robot today is all-electric — in practice servos made with Chinese rare earth materials, usually made in China — and that Boston Dynamics' hydraulic Atlas was a dead-end research prototype, which is why BD stopped it .
  • Claims China's EV companies, many also pursuing robotics, position it to scale robot production "by orders of magnitude," already orders of magnitude ahead of the US .
  • Blames US caution on culture: Americans avoid showing failures while Chinese teams push robots to breaking point on live feeds; Boston Dynamics "won't risk its image" .
  • Rebuttal: @NarrativeFilter says "China doesn't innovate, they steal" and claims the US leads by a huge margin in military robots ; teortaxesTex counters that a generation ago China had $1,000 GDP/capita and 7.6% tertiary education enrollment, so depending on foreign IP was natural, and calls the military claim "very Soviet cope" .
1) China controls near the entire supply chain for humanoid robotics, to the point that American builders have to smuggle parts out in su… [@teortaxesTex](https://x.com/teortaxesTex) To be fair, 2/3 of your argument here is supply chain, not innovation. China doesn't innovate… [@Ark__PL](https://x.com/Ark__PL) [@shaunmmaguire](https://x.com/shaunmmaguire) Exactly. We might be "behind" in civilian robots that run… [@NarrativeFilter](https://x.com/NarrativeFilter) China had $1000 GDP/capita and 7.6% tertiary education enrollment one generation ago, o… Very Soviet cope "yeah our ladas are complete trash, thankfully we have invested everything into miltech, Our Genius Scientists in Black …
AI High Signal

OpenAI researcher @kundan2510, after ~14 months at the company, says leadership fully supported work on full-duplex models (the gpt-live series) even when success seemed deeply improbable, and that this held as a general rule for OpenAI research projects; he believes very few environments could have sustained such a bet . @gdb frames this as OpenAI having built the muscle of making long-term research bets .

In my \~14 months at [@OpenAI](https://x.com/OpenAI), one of the most surprising and genuinely wonderful things about it, has been the fu… OpenAI has built the muscle of making long-term research bets: [https://x.com/kundan2510/status/2091713860528984451](https://x.com/kundan…
AI High Signal

A slow-motion video highlights a 400m champion’s running form at the World Humanoid Robot Games . @hardmaru commented that it resembles novel locomotion policies previously seen in MuJoCo simulations, except that they now work in the real world .

Slow-motion look at the 400m champion’s running form at the World Humanoid Robot Games. [![Video](https://pbs.twimg.com/amplify_video_thu… Reminds me of novel locomotion policies discovered in MuJoCo environments, except that they work in the real world! [https://x.com/erench…
AI High Signal
  • Cursor shipped Grok bot v0.18.0 with runtime source maps enabled, allowing its source code to be reconstructed; the reconstruction and downloads are now public .
  • The full T3 Code GitHub repository was accidentally left public .
The Cursor team shipped Grok bot (0.18.0) with runtime source maps enabled. Surprised nobody noticed until now. Source code reconstructed… You think that's bad? These guys accidentally left the whole T3 Code Github repo public 💀 ![](https://pbs.twimg.com/media/HQdlycebcAAJHWm…
AI High Signal

Hugging Face, the AI developer platform, is exploring a potential $13 billion sale, per Business Insider . Sharing that report, @terryyuezhuo listed Nvidia, Amazon, Microsoft, Google, IBM, Salesforce, Databricks, and Snowflake in the same post, placing them in the context of the sale .

Hugging Face, an AI developer platform, has been exploring a potential $13 billion sale, highlighting its key role in the AI ecosystem. [… nvidia amazon microsoft google ibm salesforce databricks snowflakes [https://x.com/businessinsider/status/2091602416294633477](https://x.…
AI High Signal

Comparing humanoid robotics strategies, @tszzl says Chinese demos showcase extreme kinetic locomotor skill while US efforts emphasize dexterous manipulation, reflecting US labor shortages — a divergence in the "Chinese technology tree" . @teortaxesTex counters that China has both and that flashy "runner robots" are a red herring; the real US concern is China's ecosystem of robot hands, which is more AI/compute-gated than hardware-gated .

this is an interesting area where the chinese technology tree has just diverged from american humanoids the demos we see from china invol… they have both, you just can't make an engaging Olympic Games out of competitive shirt folding. These runnerrobots with breedable hips ar…
AI High Signal

An AI system produced a Lean-verified five-line proof that Σ n!/(2n)! is irrational, addressing a question by Erdős, and the proof extends to the whole family Σ n!/((a+1)n+b)! for a≥1, 0≤b≤a . The same AI then claimed the broader affine family a≥1, b≥1-a is transcendental, using techniques beyond the author's understanding; code is on GitHub .

hello Erdős asked if Σ n!/(2n)! is irrational It is. Five-line proof. Extends to the whole family Σ n!/((a+1)n+b)! for a≥1, 0≤b≤a. Lean v…
AI High Signal

SemiAnalysis tested Anthropic's $200/mo Claude plan and found it can yield up to $8,000/mo in value, while OpenAI's $200/mo plan can yield up to $14,000/mo worth of tokens if weekly limits on long-horizon tasks are exhausted; the test was run in June .

We tested Anthropic's $200/mo Claude plan, which can yield up to $8,000/mo💰️💰️, while OpenAI's $200/month plan could yield up to $14,000/mo…
AI High Signal

Paul Graham observes that LinkedIn users are editing their work histories to add "AI" and remove "DEI" . A reply describes the trend as "brutal for everyone involved" .

LinkedIn users have been editing their work histories to add "AI" and remove "DEI". ![](https://pbs.twimg.com/media/HQbR74Xa8AA76pe.jpg) brutal for everyone involved [https://x.com/paulg/status/2091591334570397979](https://x.com/paulg/status/2091591334570397979)
AI High Signal
  • a16z Charts of the Week: humans are now the minority user of AI ; AI agents burn nearly 5x the tokens people do, up 14x since February .
  • @scottastevenson pushes back on framing: it's like saying "programs are the biggest users of computers, not humans!" — software has always been recursive .
Humans are the minority user of AI Agents burn nearly 5x the tokens people do, up 14x since February Charts of the Week: [https://www.a16… This is like saying: “programs are the biggest users of computers, not humans!” Software has always been recursive. [https://x.com/a16z/s…
AI High Signal

Speculation emerged about the provenance of "Ox Alpha": @brandonjcarl argues it is a Cursor/x.AI model post-trained on GLM, noting 'Ox Alpha = 0xA in hex = 10 = X' . @teortaxesTex disputes this, sarcastically claiming it is "a Sarvam-Google colab, on a GLM base, with Google helping on vision, Nvidia providing hardware and safety finetuning from SSI" . Both claims are unverified and part of a humorous exchange.

Ox Alpha = 0xA in hex = 10 = X Increasingly I think that this is a Cursor / [http://x.AI](http://x.AI) model that has been post-trained o… your face when you are in denial of the trivially obvious real provenance\* of ox alpha \*it's a Sarvam-Google colab, on a GLM base, with…
AI High Signal

a16z's Charts of the Week reports agents now burn nearly 5x the tokens people do, up 14x since February, concluding humans are the minority user of AI . Scott Stevenson (@scottastevenson) pushes back, calling the framing 'not intellectually honest': software has always been recursive (functions calling functions, agents calling agents) and humans remain the ultimate consumer . He adds that non-recursive AI usage should barely be a thing anymore and likens the claim to saying 'programs are the biggest users of computers, not humans' .

Humans are the minority user of AI Agents burn nearly 5x the tokens people do, up 14x since February Charts of the Week: [https://www.a16… This is not intellectually honest. Computer programs have always been recursive. Functions calling functions. Agents calling agents. Soft… Non-recursive AI usage should barely be a thing anymore This is like saying: “programs are the biggest users of computers, not humans”
AI High Signal

A viral clip shows a robot breaking a board held 2.5 m up; the poster argues height is just another coordinate in the robot's balance equation, not a risk factor, and that the feat reflects precisely timed hand speed and recoil absorption while balancing on a single point of support, not raw power . @zachtratar amplifies it as "NINJA ROBOTS," conceding such demos may be overfit to training but arguing that doesn't matter because the capability will get "even better/general in 1 year" , and notes the common discourse is centered on "yes, but does it matter?" .

It breaks a board 2.5 meters up. And it has no idea that's high, for it, that's just another coordinate. A human breaking a board at that… Imagine showing this to someone in the 90s and then telling them the common discourse was centered around "yes, but does it matter?" Guys…
AI High Signal

David Holz (Midjourney founder) is pushing for major AI labs to run every named open math conjecture through ~10,000 random agent sessions over a weekend , noting it still hasn't happened and estimating the cost at ~$150k in server rental — 8 GB300 racks for a weekend — to attempt all ~3,000 major open math conjectures .

wonder how long until one of the major labs throws every named open conjecture into 10,000 random agent sessions and just lets it rip for… this still hasn't happened? quick math suggests 8x GB300 racks for a weekend (roughly \~150k$ server cost) to do a end-run at all \~3000 …
AI High Signal

NORINCO showcased an explosion-proof humanoid robot called Fuxi at the World Robot Conference (WRC). It is designed for reconnaissance and patrol missions in high-intensity environments, assisting in emergency response. Each robot can replicate the actions of its assigned operator in real time, and the goal is to create a fleet of humanoid robots to replace human personnel in dangerous missions .

Impressive… NORINCO showcases its explosion-proof humanoid robot, called Fuxi, at the WRC. It is designed for reconnaissance and patrol m…
AI High Signal

Wan 3.0 is now live on fal, producing native 30-second clips in a single pass without stitching, with omni-reference input (text, image, audio, video, web pages, and documents) and reality-grade rendering with improved real-world motion . The model is available today through text-to-video, image-to-video, and reference-to-video endpoints .

Wan 3.0 is now live on fal. - Produces native 30-second clips in a single pass without stitching - Omni-reference input: text, image, aud… Try it here today! Text to Video [https://fal.ai/models/alibaba/wan-3.0/text-to-video](https://fal.ai/models/alibaba/wan-3.0/text-to-vide…
AI High Signal

OpenMHC was released last week to address the lack of open, large-scale wearable health research data: the biggest datasets and models are closed or heavily gated, and shared benchmarks barely exist . Commentator @iScienceLuvr calls combining wearable data with clinical data crucial for early diagnostics and continuous monitoring of patients, describing the paper as 'really important' and a first step toward that future .

Wearables are everywhere. Large-scale wearable health research data is not. The bottleneck is openness: the biggest datasets and models a… I think combining wearable data with clinical data is going to be crucially important for the develop of early diagnostics and continuous…