ZeroNoise Logo zeronoise
Post
Qwen3.8-Max Raises the Open-Weight Bar for Long-Horizon Work
23 hours ago
4 min read
463 docs
Alibaba’s 2.4T Qwen3.8-Max, MiniMax H3’s immediate serving ecosystem, and the Astra/Fable reproducibility test define the period. The brief also tracks new agent research, Japanese enterprise deployment, and the operational and compliance constraints now shaping adoption.

Top Stories

Why it matters: Frontier competition is now combining large capability claims with open access and immediate deployment economics.

Alibaba’s Qwen3.8-Max raises the open-weight ceiling. Alibaba calls Qwen3.8-Max its most capable model: a 2.4T-parameter system whose open weights, plus Qwen3.8-27B, are due next week. It claims 10+ days of autonomous coding from empty folder to production, 500+ chip-design turns, 365 days of e-commerce strategy, and vision-led self-correction. API pricing is $2/$6 per million input/output tokens ($0.25 for implicit caching); Frontend Code Arena scored it 1,668, fourth behind Opus 5 Max and Kimi K3 Max and level with Opus 5 High.

Astra’s headline is already being tested for reproducibility. OpenAI says internal Astra produced results on 10 problems open at least a decade; its roughly $2,000 figure is token cost at Sol rates, and the model formalized each argument in Lean after human manuscript preparation. OpenAI says its system generated the mathematical arguments and takes responsibility for correctness. Within 24 hours, Anthropic researcher Levent Alpöge said Fable reproduced five autonomously with a generic prompt, no internet and safeguards against leakage; only one used essentially the same argument.

Research & Innovation

Why it matters: New work is targeting the agent interface and adaptation loop—the layers that turn model capability into operational behavior.

Qwen-CUA makes the GUI a native agent interface. It sees screenshots only, with no DOM or accessibility tree, and uses mouse and keyboard across browsers, desktop apps and professional software. Qwen says it built about 40,000 verifiable tasks and rollout infrastructure with nearly 100,000 vCPUs; it reports broadly competitive results across eight computer-use benchmarks and has released the code and technical report.

SkillSmith makes skill composition an inference-time operation. Google DeepMind’s system feeds an LLM existing prefix weights plus text describing how a capability relates to a target, then emits new prefix weights. The team says this instruction-steered parametric synthesis outperforms text-only and weight-only adaptation.

Products & Launches

Why it matters: Release-day serving support is becoming part of the product, shortening the path from weights to usable applications.

MiniMax H3 pairs open weights with an inference stack. vLLM says H3 reads text, images, video and audio as one context and returns 4–15-second clips up to 2K resolution at 24 FPS with synchronized stereo audio through an OpenAI-compatible /v1/videos endpoint. It has day-zero support in vLLM-Omni and SGLang; SGLang says it matches Seedance 2.0 at one-third the cost, or can run locally without an API bill on specified GPUs.

Sakana Namazu targets Japanese enterprise workflows. Sakana launched the updated Namazu as an API, described as serving Japanese enterprises with frontier-level reasoning and built-in agentic tools. Its demonstrations cover autonomous weekly market research—planning, repeated web search, cross-checking and writing—and Japanese customer support through order-data aggregation and analysis at low unit cost.

Industry Moves

Why it matters: Deployment pressure is exposing a people-and-governance bottleneck alongside model progress.

Enterprise AI is being reorganized around operational ownership. The Turing Post reports that 95% of AI pilots show no P&L impact and identifies demand for AI Operations Leads, forward-deployed engineers, semantic modelers and evals engineers. It frames the unresolved work as securing decisions, specifying workflows, encoding meaning and verifying behavior.

Google DeepMind is adapting engineering hiring to agentic work. In its AGI Safety hiring round, all engineering interviews allow agents; Neel Nanda says candidates will work with agents all day and should be interviewed accordingly.

Policy & Regulation

Why it matters: Compliance is moving toward visible provenance requirements for model outputs.

EU transparency rules are reported to be live. A monitored update says the EU AI Act now requires models to identify themselves as AI and AI-generated images, video and audio to be labeled and watermarked; it says Anthropic, Google, Meta, Microsoft, Mistral and OpenAI have committed to comply.

Quick Takes

Why it matters: Deployment quality now depends on both harness efficiency and human checkpoints.

  • Hermes Agent: Optimizations traced through 250,000 conversations reduce turns, context load and token waste, especially for smaller and local models.
  • Codex: A user reports the app edited and published an ad, built its audience, set the budget, then stopped at the Pay button for permission while the user watched live.
  • DeepSeek V4-Flash: An OpenCode test was highly positive, but a follow-up Pi test saw the model burn 1M tokens and called it very harness-sensitive.
Qwen3.8-Max Raises the Open-Weight Bar for Long-Horizon Work
Research extraction

Yes — the linked post contains materially useful quantified and attribution detail beyond a bare announcement. It says the results were achieved by an internal version of Astra, OpenAI’s next major model; that the total tokens needed to find solutions “would cost roughly $2,000 at Sol API rates”; that human-prepared manuscripts were made with the same model; and that each argument was formalized in a Lean certificate. OpenAI also explicitly states the mathematical arguments were generated by its system while OpenAI takes responsibility for correctness.

Core framing and context

  • Published August 1, 2026.
  • Ten results are framed as solving problems open with no progress on the main result for at least a decade, in most cases much longer, spanning geometry, coding theory, circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics, with substantial interest to their communities.
  • The May AI-generated disproof of the Erdős unit-distance conjecture is cited as having inspired further developments.

The ten claimed results

  1. High-dimensional sphere packing: new upper bounds down to the Cohn–Elkies threshold.
  2. Binary and spherical codes: exponentially improved bounds on maximum size at any prescribed minimum distance.
  3. Non-sofic groups: a construction establishing existence, addressing a central open question in group theory.
  4. Connes’s rigidity conjecture: disproof that certain groups are uniquely determined by their von Neumann algebras.
  5. Arithmetic circuit complexity: new lower bounds for computing the permanent, including an arithmetic-formula lower bound of order n4/log n.
  6. Quantum parallel repetition: an exponential parallel repetition theorem for general two-player quantum games.
  7. Closest vector problem: polynomial-factor hardness of approximation, relevant to post-quantum cryptography.
  8. Ehrhart’s volume conjecture: determining, in every dimension, the maximum volume of a convex body whose centroid is its only interior lattice point.
  9. Multicolor Ramsey numbers: a superexponential lower bound, resolving Erdős problem 183.
  10. Extremal number conjectures: results on compactness and degeneracy, resolving Erdős problems 146 and 180.

Cost claim The cost figure is expressly hypothetical: “The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates.” It is tied to API token pricing and is not stated as actual spend or full research cost.

Verification, attribution, and caveats

  • Verification is described as the model’s formalization of each argument in a Lean certificate, after the arguments were prepared into manuscripts by humans using the same model; the page also says a model narration of its thinking process is released for each solution.
  • On attribution, the page argues that claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and genuine human intellectual work; OpenAI says it helped prepare the manuscripts and formalize the proofs in Lean and takes responsibility for their correctness, while the mathematical arguments themselves were generated by its system. It also acknowledges deeply held concerns, including those of signers of the Leiden declaration on AI and Mathematics.
  • The post frames community engagement as the path to placing results in context and developing the ideas further.

Limitations and gaps in this source

  • The page links to a paper and reasoning walkthroughs , but the source bundle does not include those artifacts or the Lean certificates, so the technical claims and claimed Lean verification cannot be independently checked from this material; they are self-reported in the announcement.
  • The model is described only as an “internal version of Astra,” with no model details, training, or compute specifics in this page.
  • The footnote lists follow-up work inspired by the earlier disproof but does not independently verify the ten new claims.
Ten advances in mathematics and theoretical computer science | OpenAI
AI High Signal

A new AI paper (the 'explorative-modeling-third-pretraining-axis' / XM work promoted by @AlexiGlad) is being accused of plagiarism. @KL_Div, who says they worked on IMLE for years, notes the paper's XM seems identical to IMLE and that its motivation and insights are similar . @suchenzang asks whether this amounts to 'effectively plagiarism' and claims the paper additionally plagiarized the 'mode forcing' motivation, citing a 34-page preprint "that reads like pure claude slop" instead of the 2023 paper that introduced the term . @du_yilun acknowledged it's 'not a great look' .

Was just as excited as everyone else to read this cool new paper, and it felt like a trip down memory lane! XM seems to be the same as IM… so... the explorative-modeling-third-pretraining-axis thing was effectively plagiarism? 👀 [https://x.com/kl_div/status/208405957724779755… incredible, they even plagiarized the "mode forcing" motivation, citing a 34-page preprint paper ([https://alexiglad.github.io/assets/pdf… [@du_yilun](https://x.com/du_yilun) yeah... not a great look [https://x.com/zacharylipton/status/2083759383033393465](https://x.com/zacha…
AI High Signal

On August 3, 2026, MiniMax made MiniMax-H3 publicly available on Hugging Face (https://huggingface.co/MiniMaxAI/MiniMax-H3), with a video accompanying the release announcement . AI commentator @teortaxesTex responded 'Qwen-VL carrying the industry again,' linking the MiniMax post .

MiniMax-H3 Is Now Publicly Available [https://huggingface.co/MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3) [![Video]… Qwen-VL carrying the industry again ![](https://pbs.twimg.com/media/HOw_2EGWQAAlBvn.jpg) [https://x.com/MiniMax_AI/status/208410680403287…
AI High Signal

@KL_Div says the method "XM" in a newly shared paper appears to be the same as IMLE, with similar motivation and insights, and offers context from years of working on IMLE ; their post links to @AlexiGlad's original . @cloneofsimo agrees ("100%") .

Was just as excited as everyone else to read this cool new paper, and it felt like a trip down memory lane! XM seems to be the same as IM… 100%. [https://x.com/KL_Div/status/2084059577247797554](https://x.com/KL_Div/status/2084059577247797554)
AI High Signal

Alibaba Qwen announced Qwen3.8-Max, "most capable model to date," at 2.4T parameters; open weights for Qwen3.8-Max and Qwen3.8-27B are due next week . Capabilities include autonomous coding with 10+ days of self-evolving development, long-horizon planning (500+ turns of chip design optimization, 365 days of e-commerce strategy), and native multimodal intelligence . API pricing is $2.0/M input tokens, $6.0/M output tokens, and $0.25/M for implicit caching . The model is available via Qwen Studio and API; Qwen3.8-Max and Qwen3.8-27B will also launch on Modal on day one .

📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also … Qwen3.8-Max and Qwen3.8-27B coming to [@modal](https://x.com/modal) on Day 0. [https://x.com/alibaba_qwen/status/2084100707423289643](htt…
AI High Signal

Alibaba Qwen announced Qwen3.8-Max, positioning it as 'A New Bar for Coding and Cowork' . Teknium says a new local 27B Qwen 3.8 model is coming alongside it, and that Hermes Agent was featured in the release video .

Meet Qwen3.8-Max: A New Bar for Coding and Cowork. [![Video](https://pbs.twimg.com/amplify_video_thumb/2084093069323104256/img/9MfuQy8RFs… Qwen 3.8 Max and a new local 27B Qwen 3.8 is coming! Thanks [@Alibaba_Qwen](https://x.com/Alibaba_Qwen) for featuring Hermes Agent in the…
AI High Signal

Long context compaction and training is called the most underrated job in AI labs, a small process that massively affects agent performance on high-value, multi-hour tasks, with labs differing greatly in approach . A post on Qwen 3.8 max hinted compaction worked well in its 10+ day run and paper reimplementation, though not specifically targeted . The topic is described as one of the 'unknown 6 problems' on the path to AGI .

most underrated job rn in the labs is the person handling long context compaction and training. it's a very small process that massively … The Qwen 3.8 max post hinted at compaction working well in their 10+ day run and their paper reimplementation. Although it didn't seem li…
AI High Signal

Qwen Team and XLang Lab released Qwen-CUA, a native computer-use agent that operates on screenshots only (no DOM or accessibility tree), interacting via keyboard and mouse across browsers, desktop apps, and professional software . It maintains long-horizon visual context and verifies progress; training used approximately 40K verifiable tasks and rollout infrastructure with nearly 100K vCPUs . Across eight computer-use benchmarks (everyday desktop use, long-horizon workflows, personalized computing, scientific research, web interaction, macOS, and adversarial robustness), Qwen-CUA demonstrates strong and broadly competitive capabilities, with a scaled-up Qwen-CUA-Max pushing the frontier further . Code and the technical report are open-sourced on GitHub . A separate post congratulates the team on a SOTA model release .

💻 Meet Qwen-CUA — our native computer-use agent for (almost) everything. Code, APIs, and computer use are three of the most important int… Congrats to the Qwen CUA team on releasing a sota model! [https://x.com/dunjielu1219/status/2083967342435020889](https://x.com/dunjielu12…
AI High Signal
  • Alibaba Qwen announced Qwen3.8, launching and going open-weight soon, with 2.4T parameters, claiming it is one of the most powerful models available, on par with leading frontier models and second only to Fable 5 .
  • Qwen3.8-Max-Preview is already live on Alibaba's Token Plan, Qoder, and QoderWork .
  • Simon Willison flagged that the newly named model is 'qwen3.8-max', while the earlier preview was 'qwen3.8-max-preview' ; he also asked what changed in the latest Qwen3.8 blog post (qwen3.8) versus the earlier tweet .
  • The Qwen Foundation Model Team opened an official @QwenDevs account and is holding an AMA .
Qwen3.8 is launching and going open-weight soon!🌐 With a massive 2.4T parameters, this model is continuously evolving. We believe it’s on… [@Alibaba_Qwen](https://x.com/Alibaba_Qwen) Looks like this is qwen3.8-max - that older model was qwen3.8-max-preview [@Alibaba_Qwen](https://x.com/Alibaba_Qwen) What's new in today's announcement [https://qwen.ai/blog?id=qwen3.8](https://qwen.ai/blog?id=… git init qwen\_devs README.md: Hey 👋 We're the folks from Qwen Foundation Model Team. Since we finally have an account...an AMA? ![](http…
AI High Signal
  • Alibaba unveiled Qwen3.8-Max, its most capable 2.4T-parameter model; open weights arrive next week, with Qwen3.8-27B also going open-weights .
  • Capabilities: 10+ day self-evolving autonomous coding (GitHub trace), long-horizon planning (500+ chip-design turns, 365-day e-commerce strategy), and native multimodal feedback .
  • Pricing: $2.0/M input, $6.0/M output, $0.25/M implicit caching .
  • Cline: 2% higher Terminal-Bench score than Fable 5, days after DeepSeek V4-Flash claimed similar performance; argues open weights have surpassed closed models .
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also … Qwen3.8-Max is Alibaba’s largest model yet at 2.4T params, and shows a 2% higher benchmark result on Terminal-Bench than Fable 5. This co…
AI High Signal

@goodside reports ChatGPT Work (“5.6 Sol High”) installed Blender and modeled a chess set with the board position of Byrne vs. Fischer (1956), following the capture of Fischer’s queen, balanced atop a wheelbarrow of tropical fruit in the Backrooms .

ChatGPT Work (5.6 Sol High) installs Blender and models a chess set with the board position of Byrne vs. Fischer (1956) following the cap…
AI High Signal

Qwen3.8-Max by Alibaba Qwen reached #4 on the Frontend Code Arena leaderboard with a score of 1,668, trailing only Claude Opus 5 (Max) at 1,705 pts and Kimi K3 (Max) at 1,676 pts, and on par with Claude Opus 5 (High) at 1,669 pts . It also ranks #2 in Consumer Product, #3 in Brand & Marketing, Reference-based design, Gaming, and Content Creation Tools, #4 in Data & Analytics, and #5 in Simulations . The model is priced at $2 per input MToken and $6 per output MToken, reshaping the cost-performance Pareto frontier in Frontend Code Arena; top models on the frontier include Claude-Opus-5, Kimi-K3, Qwen3.8-Max, GLM-5.2, and DeepSeek-V4-Flash .

Big news: Qwen3.8-Max by [@Alibaba_Qwen](https://x.com/Alibaba_Qwen) just landed at [#4](https://x.com/hashtag/4) on the Frontend Code Ar… Qwen3.8-Max by [@Alibaba_Qwen](https://x.com/Alibaba_Qwen) has reshaped the cost-performance Pareto frontier in Frontend Code Arena, with…
AI High Signal

In a thread recap, @KL_Div states that XM is a special case of the earliest 2018 version of IMLE, with further improvements to be shared in coming weeks .

Whew! That was probably the longest thread I've ever written on Twitter. I hope you all found that to be informative :) In summary, XM is…
AI High Signal

Alibaba's Qwen announced Qwen3.8-Max, calling it its most capable model to date, with open weights promised next week along with open-weights Qwen3.8-27B . The 2.4T-parameter model is positioned as a new bar for coding and "cowork" — 10+ days of self-evolving development from empty folder to production, system-level autonomous planning with "500+ turns of chip design optimization and 365 days of e-commerce strategy" . API pricing: $2.0/M input tokens, $6.0/M output, $0.25/M implicit caching . An early tester probing it says it has "the formatting style of sol and the reasoning laziness of opus," but is not complaining since the weights drop next week .

📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also … probed it with a math question: seems to have the formatting style of sol and the reasoning laziness of opus, but the weights are droppin…
AI High Signal

Alibaba released Qwen3.8-Max, its most capable model to date — a 2.4T-parameter multimodal model, with open weights for Qwen3.8-Max and Qwen3.8-27B coming next week .

  • Pricing: $2.0/M input tokens, $6.0/M output tokens, $0.25/M implicit caching .
  • Capabilities: autonomous coding (10+ days of self-evolving development, complete GitHub project trace), production-quality deliverables across hundreds of professions, long-horizon system-level planning (500+ turns of chip design optimization, 365 days of e-commerce strategy), and native multimodal intelligence with continuous vision feedback .
  • Arena results: #4 Frontend Code with 1,668 pts (behind Claude Opus 5 Max at 1,705 and Kimi K3 Max at 1,676; on par with Claude Opus 5 High at 1,669) ; #5 Text Arena with 1,496 pts — #1 in Medicine & Healthcare ; #2 Vision Arena with 1,305, just 13 pts behind Claude Fable 5 High .
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also … Big news: Qwen3.8-Max by [@Alibaba_Qwen](https://x.com/Alibaba_Qwen) just landed at [#4](https://x.com/hashtag/4) on the Frontend Code Ar… Qwen3.8-Max ranks [#5](https://x.com/hashtag/5) in Text Arena with 1,496 pts! In Occupational, it is: [#1](https://x.com/hashtag/1) in Me… Qwen3.8-Max ranks [#2](https://x.com/hashtag/2) in Vision Arena scoring 1,305. Second only to Claude Fable 5 (High) which has only a 13pt…
AI High Signal

DeepSeek V4-Flash 0731 is now available free in Cline; @cline calls it the first flash model they've found that performs at SOTA levels . Cline subsequently tripled the free quota for deepseek v4-flash, saying they can sustain this .

We are making the updated DeepSeek V4-Flash 0731 free in Cline. This is the first flash model we've found performs at SOTA levels, and ar… we have 3x'd the free quota for deepseek v4-flash in cline. turns out we can afford to do this quite sustainably! [https://x.com/cline/st…
AI High Signal

In a threaded exchange, @ostrisai says a model license forbids usage in the USA, EU, UK, and Korea, and even downloading the model in the USA . @teortaxesTex responds "Argentina winning!" .

Um.. If I am reading this license right, the license forbids usage in the USA, EU, UK, and Korea. We are not even allowed to download the… Argentina winning! [https://x.com/ostrisai/status/2084110556374659476](https://x.com/ostrisai/status/2084110556374659476)
AI High Signal

Alibaba released Qwen3.8-Max, its most capable model to date, with open weights expected next week along with Qwen3.8-27B . It is a 2.4T-parameter model touted as a new bar for coding and collaboration, capable of 10+ days of self-evolving autonomous coding from empty folder to production, and system-level autonomous planning with 500+ turns of chip design optimization and 365 days of e-commerce strategy; vision serves as a continuous feedback loop for planning and self-correction . API pricing is $2.0/M input tokens, $6.0/M output tokens, and $0.25/M for implicit caching . vLLM says it will support Qwen3.8-Max's open-source release on day 0 .

📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also … vLLM will support Qwen3.8-Max's open-source as usual on day-0, stay tuned🥰 [https://x.com/Alibaba_Qwen/status/2084100707423289643](https:…
AI High Signal

MiniMax released H3, an open-weights multimodal model, with native ComfyUI support on day zero . It supports five workflows: text-to-video (prompt only), image-to-video, first-and-last-frame control, reference-to-video (carrying a subject, motion, or voice), and in-place editing . H3's lead capability is multimodal context understanding — it takes images, audio, and video together and resolves them against a prompt, collapsing the five tasks into one model .

The weights are here. The nodes are native. The rest is up to your graph. 👀 H3 is in [@ComfyUI](https://x.com/ComfyUI) on Day 0!! one mod… MiniMax H3 is native in ComfyUI, day zero with the weights. Model highlights: → Text-to-video, prompt only → Image-to-video → First-and-l…