ZeroNoise Logo zeronoise
Post
MiMo-V2.6 Turns Open-Model Disclosure into a Frontier Strategy
4 min read
1090 docs
Xiaomi’s MiMo-V2.6 combines a leading open-weights benchmark result with unusually broad RL disclosure, while Grok 4.7’s launch highlights the tradeoff between capability gains and inference cost. Muse’s early distribution and commerce integration show the consumer-agent race moving into real workflows.

Top Stories

Why it matters: The frontier is now being contested on capability, deployability, and the distribution of agents into everyday workflows.

MiMo-V2.6 makes openness part of the frontier race. Xiaomi says its Pro and Flash models are omnimodal and advanced through scaled reinforcement learning; it claims Pro matches Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on Artificial Analysis, while releasing weights, a technical report, RL environments, and training code. Artificial Analysis lists Pro as a 1.02T-total/42B-active MoE at $0.13 per index task, with $0.435 per million input tokens and $0.87 per million output tokens. vLLM says both sizes have day-zero support, 1M context, multimodal inputs, FP8 weights, and seven-token speculative drafts. The strategic shift is as much inspectable infrastructure as benchmark score.

Grok 4.7 is a capability-versus-inference-cost tradeoff. ValsAI’s first evaluation placed it at 54.2%/#24, five points below Grok 4.6, although legal and medical work improved; Grok Build’s coding score rose from 47 to 56, with large gains on DeepSWE and Terminal-Bench. Artificial Analysis says the result used 81K output tokens per task versus 36K for Grok 4.6. A post-launch SDK update moved it to #10 on Vals, but a follow-up reported a score above 60% at almost three times the cost per task.

Muse is converting consumer-agent novelty into distribution. Sensor Tower data cited in a post show 264K US downloads on September 19 and 448K daily active users on September 18, ten days after launch; Appfigures estimates Muse outpaced ChatGPT over the equivalent early period. Shopify separately announced agentic checkout through Shop Pay across all Shopify stores.

Research & Innovation

Why it matters: The strongest technical signals improve the research loop—verification, collaboration, and evidence selection—not only the base model.

AI-assisted mathematics produced a concrete theorem result. A Google/CHUK collaboration reports a full proof of the Courtade–Kumar conjecture, with analytic portions verified in Lean, using human–AI collaboration and Gemini models through the Stellar Colosseum harness; the authors also say another team independently obtained a substantially different proof.

Communication is emerging as a test-time scaling axis. A paper reports that identical agents sharing a log beat independent sampling on research-heavy tasks: a five-agent Sonnet 4.6 team matched Best-of-33 on ARC-AGI-3, while four 5.6-Sol agents found a 1,957-byte MNIST model with 99.4% accuracy. The authors attribute the gains to broadcasting partial breakthroughs, while warning that communication may not help inherently serial tasks.

Question’s Gambit improves deep search before the loop begins. Its one-time clue decomposition, complementary searches, pooling, and reranking lifted GPT-5.5 on BrowseComp-Plus from 83.1% to 90.5%, with similar gains for smaller GPT-5.4-mini and DeepSeek-v4-pro; it costs 2.3–5.3 extra tool calls, and most remaining errors occur after retrieval, when evidence is used.

Products & Launches

Why it matters: Agent products are becoming executable workspaces rather than chat interfaces.

Devin Cloud in Terminal lets users create, steer, and resume sessions from the CLI, SSH into a dedicated Mac, Linux, or Windows VM, edit code, forward ports, and hand work back to a local machine.

Perplexity Computer can now generate campaign clips, product demos, and social assets with MiniMax H3 and Seedance 2.5, placing finished video beside copy and other creative in one thread for Pro and Max subscribers.

Parakeet Redux compresses NVIDIA’s 1.2GB speech model to 178MB, reportedly running at 113× real time on CPU while beating the base model on the 25-language FLEURS benchmark.

Industry Moves

Why it matters: Model competition is pulling capital toward compute independence and accelerating automation inside the labs themselves.

  • A report says DeepSeek is betting its next model on Huawei chips despite a failed Ascend 910C run, while finalizing a $7.5B round at a $75B valuation, with about $4.5B earmarked for compute.
  • The Information reports that OpenAI’s internal models now handle much of experimental model development—including GPU-kernel and execution-code optimization—and that experiments once taking years can be run in about a week.

Quick Takes

Why it matters: Small systems and serving improvements are making agent infrastructure cheaper and more controllable.

  • AutoTailor: Microsoft Research’s system reduced 1,283 generated browser APIs to 33; on 106 WebArena tasks it reached 90.6% correctness versus 87.5% for ReAct, with 57.8% lower request-token cost and 29.4% lower latency.
  • Qwen-Image 2.1: Separating fixed context from changing image positions for KV caching produced a reported 2.55× speedup.
  • Math governance: OpenAI says its independent mathematicians’ group can publish unsolicited advice and challenge the company’s decisions; members will not be paid by OpenAI.
MiMo-V2.6 Turns Open-Model Disclosure into a Frontier Strategy
AI High Signal
  • AutoJev demonstrates an autonomous model-training workflow using an agent swarm on an H200 for 20 hours at a reported total cost of $3.1K, including $1.9K for agents and $1.2K for data.
  • Built on Qwen3.8-27B, the system abandoned RL attempts in favor of supervised fine-tuning on high-quality synthetic data; the post reports that the result is competitive with or better than Jev, with improved accuracy and calibration, 260K context, multimodal input, and a Jev-compatible API/playground/inference stack reaching roughly 120 ms p95 on an H200 in an unoptimized setup. Code and model links were provided on GitHub and Hugging Face.
fun weekend project: AutoJev. i was curious to see if i could train a competitive Jev-like model completely autonomously with a swarm of …
AI High Signal
  • Muse’s “ideas” tab is being highlighted as a source of inspiration for how to use Muse; PythiaR called it a “very smart feature” and tagged Meta.
check out the muse ideas tab for inspo on how to use muse! [https://x.com/pythiar/status/2102085155091574936](https://x.com/pythiar/statu… Muse "ideas" tab very smart feature $META ![](https://pbs.twimg.com/media/HSwbQzvbUAA-WH6.jpg)
AI High Signal
  • A review of Grok 4.7 says that, compared with Grok 4.6, it is 30–80% less token-efficient, slower, more than twice as costly in real-world use, and worse on various benchmarks; the reviewer acknowledges it can still be pleasant for some engineering tasks but calls the release disappointing.
  • The review also reports unacceptably poor frontend capabilities, nonexistent 3D capabilities, and frequent Gemini-style looping.
Grok 4.5 was an incredible model for the price: fast, pleasant to use, reliable, solid default model. Grok 4.6 was a (forgivable) step in…
AI High Signal
  • PrimeBOT T1 is advertised from RMB 19,999 (approximately US$2,960; overseas pricing is still to be announced), with production reportedly scaling to one humanoid robot every 2.5 minutes and a capacity of 10,000 units per month.
  • @teortaxesTex estimates that Chinese humanoid-robot makers’ cumulative production capacity already exceeds 200,000 units annually, while arguing that current robots remain largely ineffective and that manufacturers are building capacity ahead of breakthroughs in control policies.
Meet PrimeBOT T1 🤖 Price starts at RMB 19,999 (approx. US$2,960), Overseas pricing TBA. They are scaling up and producing one humanoid ro… The cumulative capacity of humanoid robot production that I've seen from different Chinese companies seems to be >>200K/year alread…
AI High Signal
  • Theo claims GPT-6 Astra is “world class” at Blender and 3D reasoning, citing a one-shot game it created that runs entirely in a browser.
  • Theo describes Grok 4.7’s “Fishslop” run as the worst he has seen this year, noting that this was a second attempt and that the first had the submarine and fish moving backward.
GPT-6 Astra is world class at Blender and 3 dimensional reasoning. This was a 1-shot game it created, all running in browser. [![Video](h… Grok 4.7's Fishslop run is the worst I've seen this year. This is my second attempt (in the first one, the submarine and all the fish mov…
AI High Signal
  • Arav Srinivas argues that AI will become the operating system around which applications work; most non-AI apps will become headless, with social platforms a key exception because they depend on human-to-human collaboration and connection.
  • Vinod Khosla argues that “the right agent won’t be an app”: agents should meet users inside email, text messaging, and Slack, leaving any dedicated app in the background and making embedded consumer workflows strategically important.
AI is the OS. The apps will come and work around where the AI is. Most non-AI apps are going to be headless with the exception of social … “The right agent won’t be an app.” – [@vkhosla](https://x.com/vkhosla) “Where do most people live? They live in their email, they live in…
AI High Signal
  • A circulated claim says Grok 4.7 would be a 2.1T model released a few weeks later, outperform Grok 4.6 broadly while serving slightly slower and using tokens more efficiently; Theo says he has found no benchmark supporting the token-efficiency claim.
  • @teortaxesTex argues that xAI is rapidly churning and scaling base models, making its model recipes difficult to debug; they assess Grok 4.6 as an improvement over 4.5, while 4.7 appears to restart from zero.
"Grok 4.7 will be the 2.1T model released a few weeks later. This will be better than 4.6 in every way, except slightly slower to serve, … Tbh I find this a good excuse. They are churning though base models, quickly scaling; typical Elon playbook for Starships, Falcons, Optim…
AI High Signal
  • Alibaba is planning an aggressive AI scale-up: CEO Wu Yongming’s Yunqi Conference roadmap says Qwen intends to train a 5–10 trillion-parameter model, T-Head has released the Zhenwu V900, and Alibaba Cloud targets more than 20 GW of global data-center capacity by 2032.
  • Qwen is pursuing recursive self-improvement (RSI): the approach would use real-task feedback to identify weaknesses, design experiments, construct data, and continue training; next-generation models will target harder long-horizon tasks and integrated multimodal understanding and generation. Alibaba says the V900 delivers 3× the performance of the M890 and that a single cluster can scale to 500,000 cards, amid AI demand it says already exceeds supply.
  • Wu’s broader AI thesis: machine-generated thinking is currently below 3% of human output but could eventually exceed it 1,000×; he compared today’s AI coding to the 1882 electric light and argued that genuinely new machine-intelligence products have yet to emerge.
阿里巴巴 CEO 吴泳铭在云栖大会公布下一阶段 AI 路线图: Qwen 团队计划训练 5 万亿至 10 万亿参数的新模型,平头哥发布真武 V900,阿里云则计划到 2032 年将全球数据中心规模扩大到 20GW 以上。 Qwen 正在探索 RSI(递归式自我改进),让模型…
AI High Signal

Ollama reported that some requests to its deepseek-v4.1-flash model were charged at an incorrect rate over several hours, causing unexpectedly high usage consumption for some users. It reset affected users’ usage balances and restored extra amounts consumed because of the error.

Over the last few hours, some requests to the deepseek-v4.1-flash model on Ollama were charged at an incorrect rate, leading to higher us…
AI High Signal
  • Peano AI reports enabling full-parameter reinforcement learning for Xiaomi’s 310B-parameter MiMo-V2.6 model across more than 1,000 TPUs, with stable training runs exceeding 1,000 steps. The setup used vLLM for rollouts with bitwise trainer–sampler agreement; the trainer and sampler shared a TPU ICI fabric, transferring all 310B parameters in under two seconds.
We enable full-parameter RL on TPUs: MiMo-V2.6 at 310B, plus other stable training runs of 1,000+ steps across 1,000+ TPUs. With JAX, sca…
AI High Signal
  • WorldCrafter is presented as a “Consistent Video World Model with Implicit 3D-aware Memory,” with an associated paper posted on Hugging Face.
WorldCrafter Consistent Video World Model with Implicit 3D-aware Memory paper: [https://huggingface.co/papers/2609.24984](https://hugging…
AI High Signal
  • Theo gives Grok 4.7 a negative assessment: he says it is 30–80% less token-efficient than claimed, scores worse than Grok 4.6 on various benchmarks, is slower, and costs more than twice as much in real-world use—exceeding Astra’s costs.
  • He says the model remains useful for some real-world engineering tasks but criticizes its frontend capabilities, lack of 3D capability, and tendency to enter random loops.
Grok 4.5 was an incredible model for the price: fast, pleasant to use, reliable, solid default model. Grok 4.6 was a (forgivable) step in…
AI High Signal
  • Google researchers and collaborators report completing a full proof of the Courtade–Kumar conjecture, a longstanding central open problem in information theory and Boolean-function analysis; analytic portions of the proof were verified in Lean.
  • The conjecture states that, for a uniform Boolean input passed through a binary symmetric channel, coordinate projections maximize mutual information and attain the bound 1 − h(p). The work used multiple Gemini models and the Stellar Colosseum harness, while Ky and Tran reportedly produced an independent proof using a substantially different approach.
  • Stellar Colosseum is available externally through AntiGravity as a long-proof pattern, and the team plans to move its math–human collaboration features to AntiGravity as an external math collaboration platform.
I'm happy to share that we have completed a proof of the Most Informative Boolean Function conjecture (a.k.a. Courtade–Kumar conjecture) …
AI High Signal
  • Xiaomi’s MiMo-V2.6-Pro collaborated with materials researchers on metal–organic framework designs for PFAS capture, reviewing literature and patents, testing design hypotheses, and running computational “dry-lab” experiments to identify candidates for validation.
  • The team estimates the workflow delivered a 10× productivity gain, reducing the R&D cycle from one month to 2–3 days; Peking University professor Jinhu Dou said the work was comparable to that of a well-trained doctoral researcher.
From a research question to candidate materials. Xiaomi MiMo-V2.6-Pro worked with Xiaomi’s materials researchers to explore new MOF desig…
AI High Signal
  • Build a Multi-Agent System (From Scratch) added a Human-in-the-Loop chapter to Manning’s MEAP, covering a HumanInputTool, approval gates, and supervised step-by-step task execution.
  • The book’s completed Part 3 now includes chapters on building subagents—the pattern behind harnesses such as Claude Code and Codex—and A2A client/server interoperability across agent frameworks; both chapters are in editing.
Been heads down on my book for the past couple of months. Major updates to share for Build a Multi-Agent System (From Scratch) with [@Man…
AI High Signal

Qwen-Image-2.1 received Day-0 OpenVINO support from Intel, making the open-weight model ready to run optimized on Intel hardware while supporting both image generation and editing.

Day-0 OpenVINO support from [@inteldevs](https://x.com/inteldevs)! 🥳 Qwen-Image-2.1 is ready to run optimized on Intel hardware. One open…
AI High Signal

Elon Musk said Grok 4.6 was expected around August 7 as a 1.5T model with improved supervised fine-tuning and reinforcement learning, followed a few weeks later by Grok 4.7, a 2.1T model claimed to be better overall, slightly slower to serve, and more token-efficient. Theo disputed the efficiency claim, saying he had not found a single benchmark where Grok 4.7 was more token-efficient than Grok 4.6.

Interesting. Grok 4.6 releases around August 7. This will be the 1.5T model with significantly improved SFT & RL. Grok 4.7 will be the 2.… "Grok 4.7 will be the 2.1T model released a few weeks later. This will be better than 4.6 in every way, except slightly slower to serve, …
AI High Signal
  • Harvey’s gross margin reportedly dropped from about 50% to -50% by June as agent token usage increased twentyfold on rented OpenAI and Anthropic models, highlighting severe cost and unit-economics pressure in AI-agent businesses.
  • Ethan Ding interprets the shift as potentially indicating inadequate safeguards in Harvey’s enterprise contracts and that some customer agreements may have been knowingly unprofitable by design; he presents this as an implication, not a confirmed finding.
TRACKED CHANGES: Harvey’s gross margin fell from about 50% to -50% by June as agent token use spiked twentyfold on rented OpenAI and Anth… Damn this dramatically changes my priors in what the enterprises were paying for This kinda implies a wild level of CYA missing in there …