We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
🔥 TOP SIGNAL
Compare cost per completed change, not token rates alone. Anthropic says Opus 5.5 matches Claude Fable 5.1 for most tasks at 40% lower cost on typical workloads; it is now the medium-effort default in Claude Code and the Claude app for Pro, Max, and Team. OpenAI’s GPT-6 Luna starts at $0.10 per million input tokens and $0.50 per million output tokens.
The strongest coding datapoint is Boris Cherny’s internal HAProxy C-to-Rust test: Opus 5.5 and Fable 5.1 both passed nearly all tests, but Opus 5.5 finished in 9.5 hours versus 12 and cost 51% less. Treat it as a promising first-party result, not an independent bake-off.
⚡ TRY THIS
Sweep effort settings on a real task. Matthew Berman’s read of Anthropic’s FrontierCode comparison says Opus 5.5 at medium effort scored higher than max while costing under $1 versus over $5 per task. Run a representative repo task at both settings and compare test results, elapsed time, and cost before pinning a default.
Give long-running coding work an explicit end state. In Theo’s T3 Code workflow, the prompt supplied a screenshot and thread ID, asked the agent to find root causes, fix them, and file one focused PR. Say whether it should stop at the PR, babysit it until CI is green, or merge once green; Theo’s babysit instructions check only comments and CI newer than the latest push, verify bot findings, and avoid scope creep.
Try an agent-assisted formal-verification pass on risky code. Boris Cherny says a couple of short prompts with Opus 5.5 and Lean produced 16 PRs fixing bugs and race conditions in the Claude Agent SDK; he sometimes combines Lean and TLA+ to probe data flow, concurrency, and state management.
For tool discovery, keep lexical recall broad and rerank candidates with Jev. Treg takes the top 30 lexical matches on any rare query term, asks Jev whether each tool would directly accomplish—or be necessary for—the task, drops scores below 0.4, and prioritizes scores of 0.7 or higher while retaining lexical order and existing success/price adjustments. Its interleaving favored the new pipeline by 8.6:1 points excluding ties—not 9× search accuracy—and the test combines broader retrieval with Jev, so it does not isolate Jev’s contribution. Treg’s write-up.
📡 WHAT SHIPPED
Claude Opus 5.5: Anthropic lists API rates of $4/$20 per million input/output tokens and $0.20 per million cache reads—20% lower input/output prices and 60% cheaper cache reads than Opus 5. It claims 40% lower cost on typical workloads at default settings and over 30% faster output; Anthropic also cautions that benchmark margins are a less reliable guide to real-world differences than before.
GPT-6 Sol and Luna: OpenAI lists Sol at $2/$10 and Luna at $0.10/$0.50 per million input/output tokens. Its DeepSWE 1.1 results report Sol at 68.8% on max effort, within 1.1 points of Claude Fable 5 at xhigh and at about 80% lower task cost; Luna scores 66.6% at max, described as comparable to Opus 5 and Fable 5 at medium effort. These are OpenAI-reported evaluations; validate them in your own harness.
LangChain’s Patch demo turns a Slack feature request into a GitHub PR ready for review. The Managed Deep Agents recipe uses
agent.pyfor the agent/model,instructions.mdfor the repo and procedure, a Slack channel, GitHub MCP with its token in an environment-variable secret, anddefine_sandboxfor coding and tests.Simon Willison’s LLM CLI updates:
llm 0.36addsgpt-6-solandgpt-6-luna, plus a guard for single-turn models;llm-anthropic 0.29adds Opus 5.5 (llm -m claude-opus-5.5 "prompt goes here");llm-typesafe 0.1a0adds Jev support (llm install llm-typesafe).
🎬 GO DEEPER
- Matthew Berman on Opus 5.5’s effort economics: The useful segment compares effort settings and cost per task; watch it before assuming max is the best default.
- Theo on long-running Fable coding work: A firsthand PR workflow—context, root cause, review-bot feedback, and a defined stop condition.
- LangChain’s Patch walkthrough: See the Slack-to-PR agent’s setup, including its instructions file, GitHub access, and sandbox.
- Repo —
llm-typesafe: A quick way to inspect Jev’s typed yes/no, choice, and scoring calls from the LLM CLI; the release includes install and API-key setup.
Editorial take: Cheaper models widen the options; effort tuning, explicit PR completion criteria, and verification determine whether the savings turn into shippable code.
GPT‑6 Sol and Luna are listed as available in the OpenAI API under gpt-6-sol and gpt-6-luna. The launch post gives API prices per 1 million tokens: Sol $2 input / $10 output and Luna $0.10 input / $0.50 output.
- Cached input: The post says cached input-token reads receive a 90% discount. Applying that discount to the listed input rates implies $0.20 per 1 million cached input tokens for Sol and $0.01 for Luna; these effective rates are calculated from the announcement, not separately printed.
- Pricing comparison caveat: The listed GPT‑5.6 promotional rates are Sol $4/$20 and Luna $0.20/$1.20 (input/output). The post labels each GPT‑6 model “50% cheaper,” but Luna’s listed output price, $1.20 to $0.50, does not equal a 50% reduction.
- Coding: The post says GPT‑6 Sol improves substantially over GPT‑5.6 Sol on FrontierCode and matches Claude Fable 5.1 xhigh at much lower cost, without giving a numerical GPT‑5.6-to-GPT‑6 score change. On DeepSWE v1.1, max-effort Sol scores 68.8%, versus Claude Fable 5 at xhigh’s 69.9%, at approximately 80% lower cost per task; max-effort Luna scores 66.6%, described as comparable to Claude Opus 5 and Fable 5 at medium effort.
- Computer use: On OSWorld 2.0 offline, Sol at xhigh scores 60.5% versus Claude Opus 5 at medium effort’s 60.3%, at approximately 80% lower cost per task. Luna at max effort exceeds GPT‑5.6 Sol at medium effort at one tenth of its cost.
- Other GPT‑5.6 comparisons: On the post’s internal factuality evaluation, Sol makes about half as many mistakes as its predecessor; at higher effort, Luna matches GPT‑5.6 Sol at about one hundredth its cost. The post cautions that this evaluation uses conversations where users had flagged factual errors and is not representative of typical usage. Both models also show lower rates of misleading claims about their coding work than their GPT‑5.6 counterparts; the post says these evaluations deliberately test challenging situations and do not measure typical-use failure rates.
- Availability context: The post says Sol and Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, while Free and Go users can access Luna in the desktop app; the models are not yet available in Chat. ChatGPT rollout is described as gradual.
Direct answer: The supplied announcement claims strong agentic-coding results and lower typical-workload costs, publishes reduced token and cache rates, and describes developer-facing speed, security, and API changes. Its CursorBench table score and prose margin claim do not reconcile.
- Coding benchmarks: The table reports Opus 5.5 at 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode v1.1, and 57.8% on CursorBench 4.0; it shows higher scores than the listed comparison models on those coding tests. The announcement cautions that benchmark margins may be a less reliable guide to real-world differences and says its own performance gap versus Fable 5.1 is narrower than the scores suggest. Unless otherwise noted, Claude results use adaptive thinking at max effort; Terminal-Bench 4.0 uses Opus 5.5 at xhigh and GPT-6 Astra at high effort. The footnote gives a ±2.6-point standard error for Opus 5.5 on that benchmark and says the public leaderboard uses the Claude Code harness and five trials per task.
- Unreconciled CursorBench claim: The table’s 57.8% for Opus 5.5 and 41.7% for GPT-5.6 Sol imply a 16.1-point difference, whereas the coding narrative says Opus 5.5 beats GPT-5.6 Sol by 11 points; the supplied announcement does not explain the difference.
- Task cost and speed claims: Anthropic says Opus 5.5 needs less compute to serve, costs 40% less than Opus 5 on typical workloads at default settings, and generates output more than 30% faster. It attributes the task-cost reduction to both lower per-token cost and fewer tokens per task. A reported example is a 200,000-line codebase audit and fix completed in under three hours, versus over 20 hours for Opus 5, which used 2.5 times as many tokens; in an internal HAProxy C-to-Rust task, Opus 5.5 finished in 9.5 hours versus Fable 5.1’s 12 hours and cost 51% less. For per-task comparisons, the announcement says Opus 5.5 at default effort beats GPT-6 Astra on FrontierCode for roughly one-fifth the cost, matches Astra on Terminal-Bench for about 40% of the cost, and beats GPT-5.6 Sol on CursorBench by 11 points for about one-third the cost.
- Token pricing per million tokens: Opus 5.5 is listed at $4 input, $20 output, $0.20 cache reads, and $5 cache writes; the corresponding Opus 5 rates are $5, $25, $0.50, and $6.25. Fast mode is available in Claude Code and Claude Platform, offers up to 2.5× speed, and costs $8 per million input tokens and $40 per million output tokens.
- Harness and developer-facing details: The named harness detail is methodological: the Terminal-Bench 4.0 public leaderboard uses the Claude Code harness and five trials per task; Terminal-Bench-Science uses that harness and three trials per task. The announcement does not specify a new harness version or configuration. It describes a classifier that screens each action before execution, an open-source sandbox that security teams can audit, and code review intended to catch vulnerabilities before merge. It also says routine software-development bug-fixing is allowed while most cybersecurity tasks are rerouted to Opus 4.8.
- API/model behavior and availability: Preserved thinking prevents API users from editing Claude’s prior context to extract its reasoning; the announcement says it applies to Opus 5.5 and Fable 5.1 for API accounts created on or after August 31, 2026. Opus 5.5 is also no longer available with thinking switched off. The announcement lists
claude-opus-5-5as the Claude Platform model ID and says the model is available across platforms, including AWS, Google Cloud, and Microsoft Azure.
- For agentic coding comparisons, track total tokens, cost, and tool calls per prompt—not token usage per API request alone: each tool call creates another request, so request-level figures can hide higher end-to-end usage on tool-heavy tasks.
- Theo describes an informal, firsthand repo-audit evaluation: have several models find improvement opportunities in a real codebase, then use Fable 5.1 and Astra as LM judges of suggestion quality. On T3 Code, Astra led, Grok 4.7 was close, and Fable 5.1 made five suggestions versus eight for the other models, with valid suggestions but less verification. Grok 4.7’s run cost ranged from roughly peer-level to twice as much at worst, which Theo attributed to it spending extra effort verifying low-value details.
- Theo found Grok 4.7 more persistent and inquisitive, but its strengths were task-dependent: in a same-prompt, same-environment front-end/game test, he judged its result poor compared with Astra, and reports one run took nearly 90 minutes.
- In a sponsored segment, Theo says his team handles hundreds of PRs a day and now runs 8–30 checks per PR as it adds CI to verify agent work; he argues fast, inexpensive feedback matters at that scale. He describes Blacksmith as replacing GitHub Actions runners with a one-line workflow change, and says demo prompts in its Codesmith tool led him to merge three CI-improvement PRs.
- An Anthropic Claude Code team member recommends Opus 5.5 as a daily driver for general software engineering; they used it for an entire personal-site redo, while reserving Fable 5.1 for discrete planning, code review, and security tasks where the cost of a mistake matters more than cost efficiency. Opus 5.5 was available in Claude Code at launch.
- Compare cost per completed task, not token price alone: Berman relayed Anthropic’s claim that Opus 5.5 costs 40% less than Opus 5 on typical workloads at default settings, despite about a 20% reduction in listed token prices, attributing the difference to fewer tokens needed to finish tasks. In Berman’s Frontier Code comparison, Opus 5.5 at medium effort scored higher for under $1 per task than max effort at over $5; test effort settings on your own workload rather than assuming max is best.
- The Claude Code team has removed much of its system prompt and adjusted tools as models improve. Alongside the release, it introduced plugin evaluations to check whether skills and plugins are current for a new model and help update them so they do not constrain it.
- In an early-access demo, Berman’s teammate Alex asked Opus 5.5 to make a Dark Souls-like game in Unreal; Alex said it worked for about 28 hours without Slash Goal and that he now uses natural-language instructions such as “keep working until it’s perfect.” The reported result included six worlds, NPCs, an upgrade system, and roughly 15 bosses—an example of a long-running side-project workflow, not production evidence.
- Latent.Space reports Xiaomi’s open-weight MiMo-V2.6-Pro and Flash release as natively omnimodal; Xiaomi positions Pro as its most capable model and Flash as the efficiency/cost balance, and says Pro-UltraSpeed is rolling out with up to 20× faster output at the same quality. Artificial Analysis reports Pro at 1.02T total/42B active parameters, with a score of 46 on its Intelligence Index and listed rates of $0.435/M input and $0.87/M output.
- Xiaomi’s technical report describes an RL recipe using a fully asynchronous architecture, 1,568 samples per update, up to 1M context length, and 3.5–3.7B tokens per step. It mixes coding, general-agent, visual, and cyber tasks across harnesses; group-relative comparisons provide more varied reward signals for long-horizon tasks and steer training toward shorter paths and fewer tokens per task.
- For implementation, Xiaomi links coding recipes, dataset loader, and rewards and composable agent configurations. Xiaomi says the environments and training recipes will be open-sourced, but its 7K+ task datasets had not yet been released.
- A transferable orchestration pattern from the roundup is to use specialized decision models for routing, approval gates, trace scoring, tool selection, and low-cost supervision inside larger agent loops, rather than treating them as standalone agents.
-
The roundup reports Cognition introduced Devin Cloud in Terminal and
devin ssh, exposing Devin’s VM through the CLI and enabling handoff between Devin and the user’s machine; it also says GitHub Copilot had teased editable diffs and shows a Sentry-integrated canvas moving from crash report to fix. These are reported product updates/demos, not firsthand workflow accounts.
John Platt, Google Fellow and head of applied science at Google Research, described ERA as a specialized Gemini-powered research harness; the paper-era setup discussed used Gemini 2.5, and Platt said it likely would not have worked with Gemini 2.0.
- Workflow: A researcher describes a problem, and an agent helps define a scorable objective and creates a Python notebook; supplied papers can guide the initial code, while starter code is optional. ERA then mutates and tests notebook candidates in a tree and can recombine ideas from different candidates.
- Orchestration: ERA keeps shared history of attempts and results, pruning it to manage context. Its default is 10 search leaves at a time: larger batches limit cross-learning. It uses an upper-confidence-bound strategy to favor candidates with promising performance and uncertainty, rather than always choosing the current top scorer.
- Human oversight and rigor: Platt says people often need to refine the objective when the agent finds loopholes, and recommends hidden holdout sets to guard against overfitting. ERA can run for hours before a person reviews examples and steers the next round.
- Firsthand result: In Google’s contrail counterfactual modeling work, the team had been stuck for two years on estimating reflected sunlight; ERA searched confounders and found a simple model that passed synthetic-data tests the team’s earlier attempts had failed.
- Model choice (firsthand use): An Anthropic technical-staff guest says they used Opus 5.5 throughout a personal-site rebuild and sees it as a daily driver; they’d use Fable 5.1 for discrete planning, code review, or security tasks where the cost of error matters more than cost efficiency.
- Tune effort; compare cost per completed task: Matthew Berman’s recap reports Opus 5.5 at 66.4 on Terminal Bench 4.0, ahead of Astra at 57.9 and Fable 5.1 at 55.8. He says that on Frontier Code, medium effort scored higher than max while costing under $1 versus over $5 per task, and recommends experimenting with effort settings rather than defaulting to max. He also reports over 30% faster output and 40% lower cost on typical workloads at default settings versus Opus 5, arguing that cost per completed task matters more than token price alone.
- Keep the harness aligned with model capability: The Anthropic guest says the team removed much of its system prompt, changed tools, and released plug-in evaluations to check whether skills and plugins suit the newer model and help update them so they do not constrain it. For parallel-agent oversight, Berman says shorter, important-first completion summaries make it easier to switch between threads when managing 10–20 agents.
- Model choice and effort: Theo reports that Fable 5.1 is less noisy than Astra and can sustain longer runs; he defaults to High, uses X High for especially deep work, and reserves Low/Medium for deliberately limiting a rabbit hole because a failed lower-effort attempt can mean spending more tokens on a rerun. These are his firsthand usage observations.
- Prompt for a finished outcome: State where the agent should stop—such as filing a PR, monitoring it until checks are green, or merging once it is ready—and give it conditional paths when the right next step depends on what it finds. For PR babysitting, Theo’s skill says to check only comments and CI results newer than the latest push, verify bot findings against source, fix real issues, and avoid scope creep or filler comments.
- Delegate verification across agents: Theo describes Fable as the stronger code writer and Astra as a somewhat better reviewer; he has Fable ask Astra to review or test changes, then revise and retest before reporting back. For computer-use verification, he says Claude is weaker than Codex on macOS and suggests having Fable call Codex for that work.
- Gate risky changes with evidence: In a non-critical Lakebed experiment, Theo had Fable assess a risky runtime change, identify what staging lacked, and build confidence measures including shadow mode, structured refresh-failure reasons, synthetic staging traffic, and runtime health counters. He cautions that this kind of autonomous workflow is dangerous on important codebases without strong staging and QA; during the experiment, laptop-driven testing caused network problems and he told the agent to stop.
- Keep prompts and configuration lean: Theo recommends removing stale formatting rules from Agent.md/Claude.md and starting with minimal defaults, adding instructions only to address observed problems. If an agent is making unrequested changes, he relays this prompt from Anthropic’s guide: “If, while working or testing, you find pre-existing bugs, performance concerns, or behaviors the task doesn't mention, don't fix, optimize, or extend it in this change unless the requested behavior cannot work without it. Report it as a follow-up in your summary.”
LLM 0.36 added gpt-6-sol and gpt-6-luna model support. Model-plugin authors can set supports_conversation = False for single-turn models; LLM raises ConversationNotSupported if they receive assistant or tool history, and llm chat rejects them before a session starts. The first plugin using this capability is llm-typesafe—a useful compatibility guard for integrations that route requests to single-turn models.
- Simon Willison now uses GPT-6 Sol and Claude Opus 5.5 as his default models in Codex and Claude Code; he switched his Datasette Agent demo to GPT-6 Luna, which he says seems fast and competent at SQL queries and building HTML and JavaScript for Datasette Apps.
- GPT-6 Luna costs $0.10/$0.50 per million input/output tokens versus GPT-5.6 Luna at $0.20/$1.20; GPT-6 Sol costs $2/$10 versus GPT-5.6 Sol at $4/$20. Opus 5.5 costs $4/$20, 20% below earlier Opus pricing, and its cache-read price fell 60%—notable for long agent conversations, where Willison says 90%+ of input tokens are processed at cached prices.
- Caution on Opus 5.5 “max”: in Willison’s SVG-generation test, it spent so long reasoning that it hit the 128,000-token output limit without returning a response; the failure repeated, and each run cost $2.56 and took nearly 20 minutes. He suspects max can overthink to the point of breaking, though this test was not a coding-agent benchmark.
Kent C. Dodds says Kody can be useful beyond development once “everything” is wired up . The accompanying ChatGPT example chains a search for an A Christmas Carol audition email, retrieval of its Google Drive sheet music, song identification, Spotify lookup, and practice-playlist creation; it reports finding three songs and says it is building the playlist .
Simon Willison’s llm-anthropic 0.29 adds support for Claude Opus 5.5; it can be invoked through the LLM CLI with llm -m claude-opus-5.5 "prompt goes here".
Addy Osmani introduced Claude Opus 5.5, claiming it costs 40% less than Opus 5, cache reads are 60% cheaper, and it performs at the level of Claude Fable 5.1 for most tasks. He called it a step up from Opus 5 and a strong model for agentic coding and computer use. In a reply, he also said it writes more naturally and follows provided writing rules more closely than Opus 5.
@kentcdodds linked an @dabit3 post relaying a Cognition giveaway announcement: it says Devin now offers models including GPT-6 Astra/Sol/Luna, Claude Opus 5.5 and Fable 5.1, SWE-2, Gemini 3.8 Flash, Grok 4.7, Kimi K3, Inkling, DeepSeek V4.1 Flash and GLM-5.3 Flash; it also lists the Fusion Frontier harness for Fable, Astra, Sol and Opus, and cloud agents on Linux, macOS and Windows. SWE-2 was advertised as free until October 15.
Ben Tossell says Astra generated the device images and built his “50 years of devices” site; he later added a live leaderboard.
Simon Willison says GPT-6 Luna is his favorite model for building product features, citing its cost and speed; he describes it as half the price of GPT-5.6 Luna . OpenAI’s launch post describes GPT-6 Sol and Luna as faster, more affordable models carrying forward much of GPT-6 Astra’s strengths, and says their API prices are 50% below GPT-5.6 promotional pricing .
- Ben Tossell says he used Astra to generate imagery for a “50 years of devices” site and build the site itself; the project lets visitors save devices they had or wanted.
-
Ben shared a demo video made with Nilbuild’s
/video-demoskill. Nilbuild describes invoking it as/video-demo [description]; it records a web-app demo with zoom-ins and voiceover, and users can request revisions until satisfied. The skill’s repository is https://github.com/nilbuild/video-demo.
Jason Zhou says Treg compared Jev with BM25 across 3,000+ endpoints; Treg’s write-up describes Jev in its production search. For agent API/tool discovery, the team kept lexical retrieval but broadened it to the top 30 candidates matching any rare query term, then batch-scored candidates with TypeSafe Jev’s System One model using a Noul yes/no probability question: “Would calling this tool directly accomplish the task, or be a necessary step toward it, on the platform or data source the task requires?” They dropped scores below 0.4, placed scores ≥0.7 in a priority group and scores from 0.4 to below 0.7 in the next group, preserved lexical ordering within groups, and retained existing success-rate and price adjustments. Jev could judge about 30 candidates per request; the reported input cost was about $0.0002 per search at $42 per billion input tokens.
For evaluation, Treg interleaved the two result lists on the same query and credited the variant that contained or ranked higher the tool the agent actually called. The Jev variant won roughly 8.6:1 points excluding ties (rounded to 9:1)—not 9× search accuracy. A call does not establish task completion, and the comparison combines broader retrieval with Jev reranking, so it does not isolate Jev’s contribution.
- Geoffrey Huntley says Underclass now supports automatic Codex subscription resets: when a live request finds every enabled ChatGPT/Codex account unavailable, it redeems one banked reset for the account that would otherwise wait longest for natural quota recovery; the feature is enabled by default.
-
For agent secret handling, Huntley says he added Preflight as a checkpoint because agents might paste
.envinto the model; the request path isclient → preflight → underclass → model. To try it, runnix run github:ghuntley/preflight -- serveand point the harness at:8081instead of:8080; Preflight repo.
Alex Albert says he has been using Opus 5.5 with Blender and that its modeling and vision capabilities let him build an entire world from one prompt; his example is pre-earthquake Market Street, San Francisco, in 1906. The prompt first asks for a research file based on Sanborn maps, a 1906 film, historical photos, and USGS topography, recording each building’s footprint, height, facade material, occupant, sources, and confidence level. It then specifies Blender Python, reusable generators for period buildings and street objects, assembly from the sourced data, no downloaded meshes, textures, or HDRIs, and a 10-second video.
[AINews] Xiaomi MiMo-V2.6-Pro 1T-A42B: the new top Open Weights model, trained for $3M
Meet Xiaomi and other top Chinese frontier labs at AIE Shanghai (opens in new tab)!
This is a first for the “Apple of China” phone maker-turned-frontier lab: “ The MiMo-V2.6 series includes two natively omnimodal models: MiMo-V2.6-Pro is our most capable model to date, while MiMo-V2.6-Flash strikes the best balance between intelligence, efficiency, and cost. We are also rolling-out MiMo-V2.6-Pro-UltraSpeed, delivering up to 20x faster output speed at the same quality, for users who require extreme generation speed.”
[

Artificial Analysis@ArtificialAnlys
MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major …

8:11 PM · Sep 21, 2026 · 227K Views
97 Replies · 197 Reposts · 2.21K Likes
Xiaomi is not traditionally considered one of the six Chinese AI Tigers (opens in new tab), so it is very surprising to the established order of names you have come to know and love. And… it is natively omnimodal!

Xiaomi made news a few days ago when Fuli Luo, a former DeepSeek star engineer now at Xiaomi (opens in new tab), started publishing their final RL training runs (opens in new tab) live, which showed an abnormal amount of transparency in their internal metrics (opens in new tab).
As they note in their technical report (opens in new tab), they scaled RL compute along three axes:
- Larger batches and higher throughput: large batches on a fully asynchronous architecture, with 1,568 samples per update, training at up to 1M context length, and 3.5 to 3.7B tokens per step.
- More tasks and richer environments: a multi-task training suite spanning coding, general agents, visual and cyber, mixed across several harnesses so that gains in one capability reinforce the others.
- More grader compute: relative comparison within each group gives long-horizon RL tasks more precise and more diverse reward signals, closes a self-improvement loop, and steers the model toward shorter paths and fewer tokens per task.

ALL of this tooling, including the environments, will be open sourced.- the environment code and training recipes, but the complete 7k+ task datasets have not yet been released.
- Coding / software engineering: Code recipes, dataset loader and rewards (opens in new tab)
- Cyber / vulnerability reproduction: ARVO environment and training recipe (opens in new tab)
- General / knowledge work: General environment, tools and training recipe (opens in new tab)
- Visual / web development: Web-development environment and grading (opens in new tab)
- Music generation: Data preparation and music scorer (opens in new tab)
- Composable mini-harnesses: Agent configurations (opens in new tab)
- Shared environment adapters: mimoagent environments (opens in new tab)
AI News for 9/19/2026-9/21/2026. We checked 12 subreddits, 544 Twitters (opens in new tab) and no further Discords. AINews’ website (opens in new tab) lets you search all past issues. As a reminder, AINews is now a section of Latent Space (opens in new tab). You can opt in/out (opens in new tab) of email frequencies!
AI Twitter Recap
Open Models, Competition, and the China Gap
- Open models remain the central policy and market story: Nathan Lambert (opens in new tab) shared a congressional briefing on open-model performance, adoption, and U.S.-China competition, followed by a public summary (opens in new tab). The broader argument resurfaced elsewhere: @Yuchenj_UW (opens in new tab) claims frontier coding capability has plateaued since Opus 4.8, while open-source models keep closing the gap at 10–50x lower cost; @ClementDelangue (opens in new tab) similarly argues APIs are overkill for many real-world use cases and that specialized models will take share. Counterpoint: @teortaxesTex (opens in new tab) argues frontier has actually split into new higher tiers, with internal models and top closed models still well ahead.
- The release cadence from Chinese labs is now difficult to dismiss: @Thom_Wolf (opens in new tab) compiled an unusually dense ~10-week run of open releases including Kimi K3, Qwen3.8-Max, DeepSeek V4-Pro, GLM-5.3, Hy4 Preview, Atria Dawn, and more. This is reinforced by a Bloomberg-sourced note via @Polymarket (opens in new tab) that startups are increasingly building custom models on open weights to cut cost and reduce dependence on OpenAI/Anthropic. The subtext across several tweets: open-weight capability is no longer confined to midsized models; multiple teams are shipping frontier-scale MoEs with credible cost-performance stories.
Xiaomi MiMo-V2.6 and RL as the New Scaling Lever
- MiMo-V2.6 is the biggest open-model release in the set: @XiaomiMiMo (opens in new tab) launched MiMo-V2.6 Pro and Flash, described as open omnimodal models with weights, technical report, RL environments, and training code. Artificial Analysis (opens in new tab) says MiMo-V2.6-Pro debuts as the top open-weights model on its Intelligence Index (46), with 1.02T total / 42B active parameters and strong cost efficiency at $0.435/M input and $0.87/M output tokens. @victormustar (opens in new tab) notes the models are under MIT license.
- What stood out technically was not just the model, but the RL stack: @eliebakouch (opens in new tab) highlighted Xiaomi’s environment/data-factory paper for generating RL tasks from open repositories with “agents in the loop” for robustness and anti-cheating. Later commentary points to a second paper and unusually high transparency: @xeophon (opens in new tab) notes Xiaomi wants to release ~7K RL environments, and @eliebakouch (opens in new tab) emphasizes the team shipped model + tech report less than a week after the final RL run. A recurring interpretation, from @bertgodel (opens in new tab) and @Thom_Wolf (opens in new tab), is that high-quality open RL environments may now be as strategically important as pretraining corpora were in the last cycle.
- RL cost/throughput details drew attention because they compress timelines: @zephyr_z9 (opens in new tab) cites 130 hours, 75B tokens, and $2.6M for the RL run behind the result; @tianjun_zhang (opens in new tab) says the MiMo family scales RL on JAX + TPU, where scaling is “mostly a config change, not a code rewrite.” If these numbers hold up, the implication is that post-training/RL is becoming a far cheaper route to frontier-adjacent gains than many assumed.
Decision Models, Jev, and the Return of Specialized Inference
- Jev was the dominant product/theme discussion: Multiple posts converged on the same framing: this is “just” classification/routing, but with modern model intelligence and much better latency/cost. @karpathy (opens in new tab) calls it a point on the Pareto frontier for “no thinking, single token, low latency acceptable intelligence”. @willdepue (opens in new tab) describes it as a zero-shot classifier with frontier-ish intelligence, while @ClementDelangue (opens in new tab) argues the excitement shows there is large latent demand for specialized models rather than ever-larger generalists.
- The ecosystem around Jev expanded quickly: @sarah_edo (opens in new tab) built a Chrome extension that uses Jev to select and fill relevant WebMCP tools per keystroke. LangChain (opens in new tab) added Jev-as-a-judge to LangSmith; @hwchase17 (opens in new tab) and @Hacubu (opens in new tab) pushed SemIf, an open-source decision model, through the LangSmith Gateway. @omarsar0 (opens in new tab) reports using Jev to retag ~2.3K papers in 83 seconds for $0.14, with 579 high-confidence changes and manual validation of disagreements.
- The more durable takeaway is architectural: DSPyOSS (opens in new tab) argues that asking frontier agents is like managing people, while hand-writing decision-model programs is analogous to writing assembly; both extremes are useful, but brittle if overused. Several posts emphasized where these models fit best: routing, approval gates, trace scoring, tool selection, discrete document decisions, and low-cost supervision inside larger agent loops rather than as standalone “smart agents.”
Inference, Tooling, and Systems Optimizations
- Tokenizer and post-training infra both got substantive upgrades: Hugging Face’s tokenizers v1 RC (opens in new tab) claims up to 30x faster tokenization, improved multithread scaling, lower memory use, and much smaller package size; @art_zucker (opens in new tab) framed it as a new SOTA tokenization library. Separately, Halo (opens in new tab) launched as a post-training framework claiming up to 2.8x throughput over stock TRL while keeping models in native Hugging Face format.
- Inference-side engineering remains a major lever: @RisingSayak (opens in new tab) showed how KV caching is incorporated into QwenImage 2.1, separating fixed context from changing image positions and yielding a 2.55x speedup; the thread cites 50.57s → 19.86s DiT time on a warmed A100 with moderate memory overhead. vLLM (opens in new tab) published tuned serving configs for Qwen3.8-2.4T on GB300 NVL72, showing a Pareto frontier from 5K total tok/s/GPU at high throughput to 180 output tok/s/user at low latency. In video workloads, vLLM (opens in new tab) also integrated PyNvVideoCodec/NVDEC, removing CPU decode bottlenecks and reporting 2x+ throughput at 8×H100.
- Compression/quantization is still moving fast: @ZhihuFrontier (opens in new tab) summarized Tencent Hunyuan’s engineering behind packing Hy4 Preview (770B) into 214 GiB via mixed-precision quantization averaging ~2.38 bits/weight, including custom CUDA kernels in patched llama.cpp. On the edge/local side, @vikhyatk (opens in new tab) released Parakeet Redux, compressing NVIDIA’s speech model from 1.2GB to 178MB, running at 113x realtime on CPU, while beating the base model on 25-language FLEURS and staying within 0.3 WER on English.
Agents, Security, and Human-in-the-Loop Control
- Computer-use systems are becoming more productionized, but security is now central: Patrick Wardle (opens in new tab) reported a serious local-hijack flaw in Muse, arguing broad OS access makes such assistants a high-value attack surface. In contrast, DeepLearningAI (opens in new tab) highlighted Meta’s design philosophy for Muse-like agents: assume prompt injection will happen, keep real credentials away from the model, isolate tools in containers, and use an independent outbound-call gatekeeper.
- Commercial agents are also being pushed deeper into workflows: Cognition (opens in new tab) introduced Devin Cloud in Terminal and devin ssh, making the model’s VM directly accessible from the CLI and allowing handoff between Devin and the user’s machine. GitHub Copilot (opens in new tab) teased editable diffs in the desktop app, while @pierceboggan (opens in new tab) showed a Sentry-integrated canvas for moving from crash report to fix.
- A recurring systems point: inference and agent infra are shifting toward test-time compute: @sarahookr (opens in new tab) predicts compute moving from pretraining—where marginal FLOPs yield less—to test-time compute, requiring “very different infrastructure.” That theme also showed up in persistent-cache discussions for local serving, e.g. @TheZachMueller (opens in new tab) on SGLang’s multi-level hiCache (GPU/RAM/disk) for preserving KV cache across model swaps and restarts.
Top tweets (by engagement)
- Grok 4.7 release: SpaceXAI (opens in new tab) announced Grok 4.7, described as a notable improvement over 4.6 at the same price/speed. Follow-on evals were mixed: Artificial Analysis (opens in new tab) reported 56 on its Coding Agent Index with gains on DeepSWE/Terminal-Bench/SWE-Atlas-QnA, while Vals (opens in new tab) saw it rank #24 on its Vals Index, down 5 points from Grok 4.6 despite gains in legal/medical.
- OpenAI’s automated model-training workflow: A widely shared summary from @wallstengine (opens in new tab) reports that OpenAI has largely automated parts of training experimental models, including GPU kernel writing and code optimization, with internal agents collaborating and compressing some experiments from years to about a week.
- OpenAI mathematics advisory group and claims of solved open problems: OpenAI (opens in new tab) announced an independent advisory group of mathematicians to guide assessment and communication of AI advances in mathematics. Attention then shifted to the stronger claim, amplified by @AndrewCurran_ (opens in new tab) and others, that an internal OpenAI model has resolved 100+ long-standing open problems across mathematics. This was among the most consequential but least independently evaluated items in the set.
- Open-sourcing of valuable data assets: @ClementDelangue (opens in new tab) highlighted Eidon AI open-sourcing 1,274 hours of egocentric robotics data (13,451 recordings) as a rare case of a startup preserving impact for the community after shutdown.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. Qwen-Image 2.1 and Tiny Open Image Models
- Qwen-Image-2.1 released! (opens in new tab) (Activity: 2485): Qwen-Image-2.1 was released with open weights as a unified 7B image generation/editing model, positioned as a faster, lower-cost member of the Qwen-Image series (blog (opens in new tab), GitHub (opens in new tab), Hugging Face (opens in new tab)). Key technical additions include native RGBA/transparent image generation and editing, support for up to 10 reference images, multi-image inference acceleration, and localized edit control for tasks like object removal, attribute changes, product/portrait-preserving edits, panoramas, infographics, typography, and virtual try-ons. Comments primarily highlight the native transparency pipeline and local-edit interface; one example uses colored circles to target three regions simultaneously for removal, hair recoloring, and clothing replacement, suggesting interest in more controllable multi-region editing workflows.
- Qwen-Image-2.1 is reported to add native transparent image generation and transparent-image editing support, which is technically notable because alpha-channel workflows are often handled as post-processing or masking rather than directly by the image model. The linked example shows transparent-output capability: https://preview.redd.it/59fu834idoqh1.png?width=767&format=png&auto=webp&s=5fb81b135b35dac70f9d38a9995d7c1a7a2877dd (opens in new tab)
- The model appears to support multi-region local editing via visual annotations, where circled regions can be referenced in the prompt and edited simultaneously. One example asks it to “remove the metal watch in the blue circle, change the hair in the red circle to black, and replace the area in the green circle with gray short-sleeved linen pajamas,” demonstrating combined object removal, attribute modification, and region replacement in a single edit pass: https://preview.redd.it/cvh09tyvdoqh1.jpeg?width=1242&format=pjpg&auto=webp&s=32077f7420def5bec85160e2e982d6aef5efce54 (opens in new tab)
-
Several commenters highlight the model size: Qwen-Image-2.1 is described as
7Bparameters, which is significantly smaller than prior Qwen image models that commenters say were over20B. This size reduction is viewed as important for local inference feasibility, with one user specifically noting interest from the perspective of a 16GB VRAM GPU such as the RTX 5060 Ti 16GB.
- Qwen-Image-2.1 is reported to add native transparent image generation and transparent-image editing support, which is technically notable because alpha-channel workflows are often handled as post-processing or masking rather than directly by the image model. The linked example shows transparent-output capability: https://preview.redd.it/59fu834idoqh1.png?width=767&format=png&auto=webp&s=5fb81b135b35dac70f9d38a9995d7c1a7a2877dd (opens in new tab)
- Clarification on the Qwen-image-2.1 license (opens in new tab) (Activity: 948): The image is a non-meme screenshot of a Qwen Developers X post (opens in new tab) clarifying that Qwen-Image-2.1 outputs are not considered licensed “Materials”, so users retain rights to generated images/content. This matters because the model license reportedly still contains a non-commercial restriction on use of the Materials, creating ambiguity over whether commercial image generation is allowed even if generated outputs are user-owned. Commenters welcomed the clarification, with one user saying Qwen-Image-2.1 “easily beats all current Flux models.” Another noted they can run it locally via ComfyUI int8 on a
16 GB RTX 5060 Tipeaking around15.2 GBVRAM, but warned the Hugging Face LICENSE file may not yet reflect the clarified intent.-
A commenter reports running Qwen-Image-2.1 locally in ComfyUI using
int8quantization on a 16 GB RTX 5060 Ti, with VRAM peaking around15.2 GB. They describe the model as suitable for local testing but note that licensing uncertainty around generated outputs was the main blocker for broader/client use.-
Several commenters highlight a legal/implementation mismatch: the Hugging Face README was apparently clarified, but the actual LICENSE file still contains Section
2(b)language prohibiting commercial “use” of the Materials. One user emailed[email protected]asking whether the license text will be updated, because the tweet/README intent may not be sufficient for client or commercial work. - The key technical/legal distinction being debated is whether “commercial use not allowed” applies only to serving, redistributing, or monetizing the model/materials, versus also restricting outputs generated by the model. Commenters argue that until the canonical license file is updated, downstream users comparing it with permissive Apache-2.0/MIT-style model licenses may reasonably avoid commercial workflows despite the clarification.
-
Several commenters highlight a legal/implementation mismatch: the Hugging Face README was apparently clarified, but the actual LICENSE file still contains Section
-
A commenter reports running Qwen-Image-2.1 locally in ComfyUI using
Keep reading with a 7-day free trial
Subscribe to Latent.Space to keep reading this post and get 7 days of free access to the full post archives.
- Latent.Space reports Xiaomi’s open-weight MiMo-V2.6-Pro and Flash release as natively omnimodal; Xiaomi positions Pro as its most capable model and Flash as the efficiency/cost balance, and says Pro-UltraSpeed is rolling out with up to 20× faster output at the same quality. Artificial Analysis reports Pro at 1.02T total/42B active parameters, with a score of 46 on its Intelligence Index and listed rates of $0.435/M input and $0.87/M output.
- Xiaomi’s technical report describes an RL recipe using a fully asynchronous architecture, 1,568 samples per update, up to 1M context length, and 3.5–3.7B tokens per step. It mixes coding, general-agent, visual, and cyber tasks across harnesses; group-relative comparisons provide more varied reward signals for long-horizon tasks and steer training toward shorter paths and fewer tokens per task.
- For implementation, Xiaomi links coding recipes, dataset loader, and rewards and composable agent configurations. Xiaomi says the environments and training recipes will be open-sourced, but its 7K+ task datasets had not yet been released.
- A transferable orchestration pattern from the roundup is to use specialized decision models for routing, approval gates, trace scoring, tool selection, and low-cost supervision inside larger agent loops, rather than treating them as standalone agents.
-
The roundup reports Cognition introduced Devin Cloud in Terminal and
devin ssh, exposing Devin’s VM through the CLI and enabling handoff between Devin and the user’s machine; it also says GitHub Copilot had teased editable diffs and shows a Sentry-integrated canvas moving from crash report to fix. These are reported product updates/demos, not firsthand workflow accounts.