We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Top Stories
Why it matters: Frontier competition is now combining large capability claims with open access and immediate deployment economics.
Alibaba’s Qwen3.8-Max raises the open-weight ceiling. Alibaba calls Qwen3.8-Max its most capable model: a 2.4T-parameter system whose open weights, plus Qwen3.8-27B, are due next week. It claims 10+ days of autonomous coding from empty folder to production, 500+ chip-design turns, 365 days of e-commerce strategy, and vision-led self-correction. API pricing is $2/$6 per million input/output tokens ($0.25 for implicit caching); Frontend Code Arena scored it 1,668, fourth behind Opus 5 Max and Kimi K3 Max and level with Opus 5 High.
Astra’s headline is already being tested for reproducibility. OpenAI says internal Astra produced results on 10 problems open at least a decade; its roughly $2,000 figure is token cost at Sol rates, and the model formalized each argument in Lean after human manuscript preparation. OpenAI says its system generated the mathematical arguments and takes responsibility for correctness. Within 24 hours, Anthropic researcher Levent Alpöge said Fable reproduced five autonomously with a generic prompt, no internet and safeguards against leakage; only one used essentially the same argument.
Research & Innovation
Why it matters: New work is targeting the agent interface and adaptation loop—the layers that turn model capability into operational behavior.
Qwen-CUA makes the GUI a native agent interface. It sees screenshots only, with no DOM or accessibility tree, and uses mouse and keyboard across browsers, desktop apps and professional software. Qwen says it built about 40,000 verifiable tasks and rollout infrastructure with nearly 100,000 vCPUs; it reports broadly competitive results across eight computer-use benchmarks and has released the code and technical report.
SkillSmith makes skill composition an inference-time operation. Google DeepMind’s system feeds an LLM existing prefix weights plus text describing how a capability relates to a target, then emits new prefix weights. The team says this instruction-steered parametric synthesis outperforms text-only and weight-only adaptation.
Products & Launches
Why it matters: Release-day serving support is becoming part of the product, shortening the path from weights to usable applications.
MiniMax H3 pairs open weights with an inference stack. vLLM says H3 reads text, images, video and audio as one context and returns 4–15-second clips up to 2K resolution at 24 FPS with synchronized stereo audio through an OpenAI-compatible /v1/videos endpoint. It has day-zero support in vLLM-Omni and SGLang; SGLang says it matches Seedance 2.0 at one-third the cost, or can run locally without an API bill on specified GPUs.
Sakana Namazu targets Japanese enterprise workflows. Sakana launched the updated Namazu as an API, described as serving Japanese enterprises with frontier-level reasoning and built-in agentic tools. Its demonstrations cover autonomous weekly market research—planning, repeated web search, cross-checking and writing—and Japanese customer support through order-data aggregation and analysis at low unit cost.
Industry Moves
Why it matters: Deployment pressure is exposing a people-and-governance bottleneck alongside model progress.
Enterprise AI is being reorganized around operational ownership. The Turing Post reports that 95% of AI pilots show no P&L impact and identifies demand for AI Operations Leads, forward-deployed engineers, semantic modelers and evals engineers. It frames the unresolved work as securing decisions, specifying workflows, encoding meaning and verifying behavior.
Google DeepMind is adapting engineering hiring to agentic work. In its AGI Safety hiring round, all engineering interviews allow agents; Neel Nanda says candidates will work with agents all day and should be interviewed accordingly.
Policy & Regulation
Why it matters: Compliance is moving toward visible provenance requirements for model outputs.
EU transparency rules are reported to be live. A monitored update says the EU AI Act now requires models to identify themselves as AI and AI-generated images, video and audio to be labeled and watermarked; it says Anthropic, Google, Meta, Microsoft, Mistral and OpenAI have committed to comply.
Quick Takes
Why it matters: Deployment quality now depends on both harness efficiency and human checkpoints.
- Hermes Agent: Optimizations traced through 250,000 conversations reduce turns, context load and token waste, especially for smaller and local models.
- Codex: A user reports the app edited and published an ad, built its audience, set the budget, then stopped at the Pay button for permission while the user watched live.
- DeepSeek V4-Flash: An OpenCode test was highly positive, but a follow-up Pi test saw the model burn 1M tokens and called it very harness-sensitive.