We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Routing inside the harness: 64% cheaper, same merge rate
LangChain wrote up how it built a model router for Open SWE, its internal coding agent, and published the results. It picked three tiers from different providers: GLM-5.3-Flash (xhigh) as fast, GPT-5.6 Sol (medium) as balanced and GPT-6 Astra (low) as performance . The router middleware runs once, on the thread's first human message, and has three parts :
- a base prompt telling the classifier to pick the least expensive model likely to finish the task
- short, plain-language criteria for each tier
- a classifier model
The tier criteria come from two sources: LangChain's own analysis of its task mix, and each provider's prompt guide. Because those criteria depend so much on the agent's own tasks, the router lives in the harness rather than a generic gateway . The first classifier was an LLM with structured output. Switching it to the Jev decision model made classification almost 50× faster. The chosen model then handles the whole thread .
They measured outcomes with merged PRs per thread and thumbs up/down ratings. Merged PRs were the stronger signal because few threads got rated . Results from a 973-thread A/B test against always using Astra:
- Quality: 29.2% of routed threads ended in a merged PR, against 27.3% for control (p=0.49) .
- Cost: the median thread fell from $2.61 to $0.94, a 64% drop. The mean fell 42% and p90 fell 37% .
- Routing mix: 56% of threads went to balanced, 34% to fast and 10% to performance. Median cost per thread ranged from $0.097 on fast to $2.88 on performance .
The opposite test, router against always using the fast model, was stopped within a day. Engineers said the low quality was disrupting their work . In short, routing saves money, but sending everything to the cheapest model doesn't work.
Related: LangChain's new smithtune CLI turns an agent's own trajectories into fine-tuning data . LangChain says OpenSWE Review's precision rose from 62.9% to 81.5% with 29% fewer tool calls per review .
Hand off verification too, not just coding
In a new video, Theo estimates that only 10–15% of his tokens go to writing code. The rest goes to verifying it, so the PR is actually done by the time he looks . He also says that if undoing a bad change takes more than 15 seconds, you should fix your systems before vibe coding . Habits from the video:
- Give the agent the problem, not your solution. He now sends a bug screenshot with "Fix it." If the agent fails, that tells him the problem deserves his own thinking .
- Ask for everything up front: build it, file a PR, spin up a Tailscale environment for him to try, then babysit the PR . If you state the conditions under which the agent may merge, the thread can close without you ever opening it .
- Treat threads as a to-do list. He fires prompts into the background (Cmd+Enter in T3 Code), only opens threads marked done or needing input, and spreads work across his Linux boxes, with auto load-balancing available .
On exploratory work, Theo argues that the branch you explore on and the PR you merge can, and often should, be built differently: "assume all PRs will be closed, not merged" . Geoffrey Huntley goes further on review. Read the properties your tests express, and if performance is a problem, "run a loop and optimise" while those properties keep holding . The things he says you must read are public SDK APIs, your verification and agent traces .
The Pragmatic Engineer quotes DHH saying 37signals is "done writing code by hand." When it happens, it's treated like a bug: you ask why the agent couldn't produce the result, then fix the "factory" . 37signals is now building native mobile apps, moving backend services from Ruby to Rust because agents write good-enough Rust, and keeping Rails because its conventions suit agents .
Claude Code mods
Claude Code now supports mods, which can change its behavior, customize the UI or swap in your own features. You can write one in a few lines of TypeScript or have Claude build it. Mods ship inside plugins and install with /plugin in the CLI or desktop app . Boris Cherny's pitch is that everyone works differently, so build your own and share them as plugins .
Dots: open autonomous PRs as drafts
Alexander Embiricos described an early internal incident. A dot watching a feedback channel found a bug and opened a non-draft PR. It had inferred the engineer's habit of setting auto-merge, and the PR merged once a colleague approved it. "Dot should set PRs in draft mode," he said . If you give an always-on agent repo access, make draft PRs the default.
Riley Brown's notes after 24 hours with a dot:
- Voice is the standout, but he hit problems in sessions of 30 minutes or more .
- He wants a better view of the Codex threads the dot spins up .
- He found it confusing to choose between the VM browser, the Codex browser and Chrome .
On capacity, Tibo says GPT-6.1 Sol is back to normal speed after a load spike, and a global usage reset for paid ChatGPT accounts lands at 10am PST on the day after his post . ThePrimeagen says about 60 minutes of prompting on a "$500 plan" used 56% of his weekly limit .
Smaller items
- Cursor added GLM 5.3 and GLM 5.3 Flash. Cursor says GLM 5.3 Max is the best-scoring open-weight model on CursorBench 4.0 .
@langchain/mcp-adapters2.0 supports the latest stateless MCP, adds tools that can check in with users and makes auth easier .- Pi 1.0 shipped with Pi Durable (post) .
- Persistent subagents: in a Latent Space interview, swyx calls ephemeral Codex subagents his biggest pain and Alex Zhang agrees that writing state to the filesystem is "the big trick." Zhang's Pi-based Prime Agent offers subagents that outlive the root agent and can be prompted later .
- Figma MCP workaround: Kent C. Dodds suggests forking his Kody Figma setup and asking your agent to set it up, which he says takes about five minutes . He builds Kody.
- Btrfs: Peter Steinberger found copy-on-write great for worktrees but terrible for SQLite. The next OpenClaw update moves the database to a NOCOW location .
- Making agent output easier to understand: Karpathy suggests asking for ASD-STE100 controlled English (or "80% of the way" to it), diagrams, HTML pages, or "a 3b1b style video explainer on X" .
- Security: Matthew Green notes that agents in separate sandboxes left instructions for each other in a shared package cache, and those instructions changed what the recipients did. Swap the cache for email, Slack or shared docs and you have the makings of an agent worm .
- Before a native iOS build, create a fresh Claude Code project folder and ask the agent to build and run a Swift “Hello World” in Xcode Simulator, verifying the setup before starting the app; the walkthrough required Xcode 27.1 beta for the Duo simulator.
- For the prototype, use Opus 5.5 at high effort with a product brief and visual references; start with local storage, then add image search with SER API and background removal with remove.bg. The walkthrough later added Convex through its Claude Code plugin for real-time storage and share links, and demonstrated live updates on a shared, view-only board. The presenter pasted API keys into the prompt while acknowledging that this is not generally best practice.
- Test in the simulator and give the agent a grouped list of specific design and behavior changes—for example, decluttering controls, animating sticker creation, saving stickers under their search terms, and filtering existing stickers while typing; the walkthrough showed the animation and filtering after revision.
- The initial build took about 37 minutes; the database extension is reported as taking roughly 50–57 minutes.
- Recent GPU Mode leaderboard solutions were mostly AI-generated, but an expert member’s AI-assisted, directionally guided kernel was described as essentially the only top-10 solution stable in end-to-end systems; the speaker also flags kernel verification and reward-hacking risks. A useful review discipline is to diagram the code and keep working until every part is understood; domain expertise helps steer and verify the model rather than relying on more compute alone.
- For persistent coding-agent work, externalize state to the filesystem rather than relying on ephemeral subagents; the discussion presents this as a workaround for short-lived Codex subagents. Prime Agent demonstrates a code-centric alternative: it is Pi-based, with Python/IPython as its only direct tool and other tools callable as code modules or scripts; code can coordinate agents, persistent subagents can be revisited and prompted, and its continual harness can let the agent modify skills, available subagents, and its system prompt.
- A suggested latency optimization is to infer likely tool calls from code as it is being written and launch them before the code is finished, rather than waiting for sequential calls; the discussion says functional languages support this pipelining and names Effect TS as a TypeScript option.
- For long-context coding-agent work, an RLM-style harness keeps context in code-managed memory and lets the model use code to call tools and recursively invoke itself or subagents, rather than relying only on an ever-growing prompt trajectory. Prime Agent is a concrete implementation built on Pi: IPython is its only exposed tool, while other tools are loaded as Python modules or Bash scripts; its continual-harness feature can modify skills, available subagents, and the system prompt, and persistent subagents can outlive the root agent’s normal runtime and be revisited. Externalizing state to a filesystem is highlighted as a way to preserve work when subagents would otherwise be ephemeral.
- Zhang says RLM strategies learned on short tasks transferred to tasks 8–30× longer and across different task types, because the high-level solution can remain similar while subtasks differ. He contrasts this with compaction, which can be faster and cheaper, and estimates that roughly 95% of some swarm exploration may be useless search.
- To reduce tool latency in code-driven agents, start predictable tool calls while the model is writing code—or as soon as enough code exists to infer the calls—instead of waiting for sequential execution; static analysis or JIT compilation may help identify calls early.
- For AI-generated GPU kernels, verify end-to-end stability rather than trusting leaderboard results: Zhang says reward hacking is a known issue and that the only top-10 solution he describes as stable in end-to-end systems was an expert’s AI-assisted, human-steered kernel. He argues that domain expertise makes people stronger verifiers and can reveal solutions that avoid enormous brute-force token search.
- On GPU-kernel work, Alex Zhang diagrams his code and works until he understands each part; he sees domain knowledge as leverage for steering and verifying AI, not merely spending more compute. Recent GPU Mode leaderboard solutions were largely AI-generated, but he observed only one top-10 kernel that was stable in end-to-end systems, and flagged verification failures and reward hacking as persistent issues.
- Prime Agent, built on Pi, gives the model code as its only direct tool, with other tools exposed through Python modules or bash scripts. Its subagents communicate through code and can persist beyond the root agent’s runtime, so they can be revisited and prompted further; for ephemeral subagents, Alex’s practical workaround is to externalize state to the filesystem.
- A speculative-execution idea for coding agents: use static analysis of code being written to launch likely tool calls in advance, rather than waiting for sequential calls; the discussion notes this is easier to pipeline in functional languages and mentions Effect TS for TypeScript.
- Start with the problem, not a prescribed fix: bring the agent in while framing the issue, ask whether there is a simpler solution, or send a bug screenshot with “Fix it”; if it fails, then invest more human effort in diagnosing and steering.
- Run threads asynchronously as a task inbox: launch work in the background, move on to other tasks, check threads when they are done or need input, and settle completed ones to keep the queue focused.
- Ask for merge-ready work, not just code: Theo’s example prompt includes building the feature, creating an environment to try it, filing a PR, and monitoring the PR; specify conditions under which the agent may merge. He estimates coding takes only about 10–15% of his token use, with most spent verifying, and recommends making bad changes quick to revert.
- Scale across machines for longer-running work: T3 Code lets him choose a connected machine or auto-load-balance threads, and he recommends using a spare Linux computer for work that should continue while he is away from his laptop.
- In the video's account of OpenAI's keynote, Alfred was asked to remove an old inventory API and reportedly traced dependencies, updated integrations, ran tests, and opened three PRs . The narrator says the subsequent live demo mainly showed that live demos are still difficult, so treat this as a demo workflow rather than proof of reliable operation .
- The described Decisions API pattern gives a small model a finite list of allowed answers and gets one back in a few hundred milliseconds, avoiding free-form JSON parsing and retry loops .
- In a sponsored segment, the creator says he added Fastino's skill to OpenCode in about five minutes; it lets his agent route developer requests to a coding model or internal tool and extract repository file paths and function names . Fastino claims its open-weight Gliner model can run these decisions locally up to 8× faster than Jev, and that Glide performs better than Jev on intent routing, fact-checking, and hallucination checks .
- At the time of the video, Gemini 4 Argon had been announced but not released; Sanfilippo says Google was pursuing cybersecurity checks before a controlled release. Google reportedly used it internally to migrate codebases to Rust, and its claimed output limit was 1 million tokens, up from 64K—a scale Sanfilippo connects to generating large amounts of code in one run.
- Sanfilippo reports Argon pricing of about $10 per million output tokens when uncached, with a 95% reduction when cached; he compares that with roughly $50 elsewhere, so treat these as his reported figures.
- Practical advice: avoid annual or organizational lock-in to one provider; switch as model quality and token costs change to keep access to the best model at the lowest cost.
- He had heard Google Antigravity’s harness works well, but explicitly says he had not tested it much recently; this is secondhand rather than a hands-on recommendation.
-
Codex’s
/gofeature lets the model pursue a problem for a long time; internal trials used long-running agent threads, and the team chose a cloud-hosted, embodied agent for harder, extended tasks such as “babysit this PR.” -
Embiricos uses the same Dot context across Slack threads, forwarded DMs, and the app. In a Space page, he tags it to flesh out a draft, pull data, contact a teammate, and check off completed work; page-level agent instructions, analogous to a repo
AGENTS.md, can direct it to keep the page updated. - In one internal workflow, a Dot monitoring a feedback channel found a bug and opened a PR, inferring the engineer’s preference for direct PRs and auto-merge; the PR was not a draft and merged after another engineer approved it. Embiricos says it should have been opened as a draft, a concrete caution for autonomous PR workflows.
- For developers running models locally, decoding speed alone can mislead: Sanfilippo says long thinking can make even a model generating 100 tokens/second feel slow, while producing a large reasoning trace at 15–20 tokens/second is especially impractical. He cites DeepSeek prefill rates of roughly 500–1,000 tokens/second on DGX Spark, depending on version and quantization, as a reason to watch reasoning-token volume as well as throughput.
- He points to Astra as an example of a model that produces much less output and thinks less, and predicts open-weight models will increasingly be optimized for low-token reasoning. If models become more capable per token, he argues, current local hardware such as DGX Spark and Strix Halo could become more useful rather than automatically requiring an upgrade; this is a forecast, not a reported coding-agent workflow.
Karpathy suggests asking an LLM to present information in a more usable form: use ASD-STE100 for cleaner writing, or request “80% of the way” to that standard for less rigid output; ask for a diagram or an interactive HTML page when those formats make the result easier to understand. For a custom explainer video, he suggests prompting “Create a 3b1b style video explainer on X” and supplying an ElevenLabs API key, or asking the LLM to find free alternatives that run on local compute. His broader workflow idea is to use abundant model intelligence and code to create large, bespoke, discardable artifacts—such as web apps or video explainers—so human effort can shift toward oversight and understanding.
Matthew Green warns that sandbox isolation may not prevent agent-to-agent payload propagation: sandboxed agents left instructions in a shared package cache that changed recipients’ behavior, and shared channels such as email, Slack, documents, or WhatsApp could provide similar propagation paths between personal agents such as Muse—ingredients for a worm, in his assessment.
Claude.dev launched as a developer hub with Claude Code and API guides, engineering deep dives, and tips from the teams building Claude . Its live material includes how the Claude web app was made 3× faster, advice on using Opus 5.5 effectively, and automating evaluation hillclimbing with Claude .
Organizations should give people beyond designated software engineers the ability to develop software; the post argues that coding has been commoditized while corporate access has not, and urges business leaders to prioritize broader software contribution. Engineers’ role, in this model, is to design systems and safety controls that let everyone ship to production safely.
- Google reports Gemini 4 Argon agents freed more than 300 TiB of data-center memory and are migrating over 800,000 lines of C/C++ kernel code to Rust. In one video-decoder example, agents replaced 32,000 lines of SIMD code with safe Rust; the existing Rust port became 2.7× faster with identical output. Argon access is limited to government users and trusted cyber defenders in the Fairwind Program while Google refines guardrails before broader developer access.
- Context Language Models offer a different agent-context design: treat context as an editable file rather than an append-only log, with context-management policies learned in the model instead of an external harness. The reported result was 65% higher performance at the same compute on a 24-hour multi-repository agent-swarm task.
- OpenAI’s DevDay agent stack includes dots—persistent agents with their own cloud computers—a Decisions API, and computer use; ChatGPT Sites can host MCP servers as installable plugins.
- A speed caveat for agent workflows: an OpenAI hands-on report found roughly 8× faster generation produced only 2–4× faster end-to-end agent tasks because tool latency dominates; gains were largest for computer use, where UI actions respond in milliseconds. Cloudflare’s agent Containers report a 648 ms p50 time-to-interactive (6× faster), with snapshots in beta; its AutoRouter showed about 30% lower spend in internal tests.
ThePrimeagen reported that after buying a $500 plan and prompting for about 60 minutes, he had used 56% of his weekly limits . In a follow-up, he wrote that he understood why people hate “astra,” without giving further detail .
Ben Tossell shared a Factory link for $200 in credits . He later said bots had ruined the fun and that he would run another credit giveaway another time .
- Brown’s setup pattern was to ask Claude Code to get the simulator working, then build and run a Swift Hello World app before starting the real project; his iPhone Duo simulator setup required Xcode 27.1 beta .
- For the app, he used Opus 5.5 at the “high” setting and supplied a device-specific brief plus reference images; he began with local storage, then separately asked Claude Code to add Convex for cloud data and share links. Image search and background removal used SER API and remove.bg; the initial app generation took about 37 minutes .
- He tested changes in the simulator by sending a grouped list of concrete UX requests, including fewer buttons, a sticker-conversion animation, and saving/searching existing stickers by their search terms. The finished share view was read-only and updated in real time .
- Credential-handling caveat: he pasted API keys into Claude while acknowledging that doing so is not generally best practice .
A user says kody.codes, linked to Kent C. Dodds, has become part of their daily workflow and calls it “incredible”; Kent reacted with “Holy smokes,” but the posts do not explain what the site does or how it is used.
Kent C. Dodds points to celld as how he makes Kody self-hostable. The linked author describes celld as a distributed runtime for scalable applications with S3 as its only dependency, and argues that traditional JavaScript runtimes are needed for builds and scripting, not serving applications.
ThePrimeagen says a $500 plan and about 60 minutes of prompting consumed 56% of his weekly limits; he does not identify the plan or what he was prompting.
[AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1M output
GDM last shipped a larger-than-Flash model in February (3.1 Pro (opens in new tab)), and after successive incremental 3.x Flash versions and the big GDM management shakeup (opens in new tab) last month, the largest question for GDM was when they would catch up to peers who have in the meantime launched Fable and Astra class models.
Well, Argon’s here (opens in new tab), with VERY respectable benchmarks (SOTA in 13 of 19 credible benchmarks)… but only accessible in limited cybersecurity preview, though access is promised “as soon as possible (opens in new tab)”:
We like the experimental Long Decode Continuation (opens in new tab), which increases output tokens up to 1M as an industry first.
AI News for 9/29/2026-9/30/2026. We checked 12 subreddits, 544 Twitters (opens in new tab) and no further Discords. AINews’ website (opens in new tab) lets you search all past issues. As a reminder, AINews is now a section of Latent Space (opens in new tab). You can opt in/out (opens in new tab) of email frequencies!
AI Twitter Recap
Gemini 4 Argon: Google Returns to the Frontier
Launch: Google DeepMind introduced Gemini 4 Argon for coding, enterprise knowledge work and cyber defense (@GoogleDeepMind (opens in new tab), @sundarpichai (opens in new tab)).
Availability: Access starts with government users and trusted cyber defenders in the Fairwind Program. Google says it will refine guardrails before opening access to developers, enterprises and consumers (@Google (opens in new tab), @demishassabis (opens in new tab)).
Output limit: Google cites an industry-leading 1M-token output limit, up from 64K (@GoogleAI (opens in new tab), @TheRundownAI (opens in new tab)).
- Measurement note: Vals lists 262K max output. Artificial Analysis reached 1M output tokens through Long Decode Continuation, a new API feature that pauses long responses and resumes them across calls (@ValsAI (opens in new tab), @ArtificialAnlys (opens in new tab)).
Pricing: Standard pricing is \$4/\$20 per 1M input/output tokens. A 50% introductory discount brings it to \$2/\$10, with no end date announced. Cached input gets a 95% discount (@_philschmid (opens in new tab), @ArtificialAnlys (opens in new tab)).
Google’s claimed results: Argon takes first place on 13 of 19 published benchmarks against GPT-6 Astra and Claude Opus 5.5. On DeepSWE it scores 77.9%, versus 74.2% for Opus 5.5 and 74.1% for Astra (@TheRundownAI (opens in new tab)).
Internal deployments: Google reports that Argon agents freed more than 300 TiB of data-center memory and are migrating more than 800K lines of C/C++ kernel code to Rust (@kimmonismus (opens in new tab)).
- Video decoder: Agents replaced 32K lines of SIMD code with safe Rust, making the existing Rust port 2.7x faster with identical output.
Research use: The team says internal agent loops built on Argon helped complete the CK conjecture (@mirrokni (opens in new tab)).
Artificial Analysis evaluation: Argon scores 53 on the Intelligence Index, matching GPT-6 Astra (53) and edging GPT-6.1 Sol (52) (@ArtificialAnlys (opens in new tab)).
Cost per task: At discounted pricing it costs \$1.99 per task versus \$3.26 for Astra; standard pricing would raise this to \$3.98.
- Token use: The savings come from price, not efficiency. Argon averages 62K output tokens per task against Astra’s 27K.
Agentic work: It ranks #1 on AutomationBench-AA at 77.5% and scores 57% on Terminal Bench 4, behind Sonnet 5.5, Opus 5.5 and Astra.
Hallucination: Its 15% rate on AA-Omniscience compares with 51% for Astra. The tradeoff is lower accuracy: 50% versus Astra’s 63% (@aipulseda1ly (opens in new tab)).
Vals evaluation: Argon is #1 on the Vals Index at 68.9%, at an average \$15.68 per task (@ValsAI (opens in new tab), @ValsAI (opens in new tab)).
Coding: It built 30 Vibe Code Bench apps perfectly, against 25 for Opus 5 and 24 for Astra (@ValsAI (opens in new tab)).
Terminal and security: Terminal-Bench 4.0 rose from 19.0% to 57.6%. It scores 70% on CyberBench proof-of-concept tasks and 100% on IOI 2024–2026 (@ValsAI (opens in new tab)).
Efficiency: It uses about a quarter of Sonnet 5.5’s output tokens on Vals Index tasks (@ValsAI (opens in new tab)).
Arena and other evals: Argon is #1 in Text Arena at 1525 and #8 in Code Arena WebDev at 1679 (@arena (opens in new tab)).
Agent Arena: It ranks #8 overall and #1 for steerability on a preliminary 3K sessions (@arena (opens in new tab)).
PostTrainBench: It scores 45.3%, up from 21.99% for Gemini 3.1 Pro (@karinanguyen (opens in new tab)).
Skepticism: Some observers questioned the published numbers.
Legal benchmark: Argon’s reported 19.6% on Harvey’s legal benchmark trails Muse Spark 1.2’s listed 25.42% (@BlackHC (opens in new tab)).
Other critiques: Commentators raised possible preference-data benchmaxxing and objected to some figures, including DeepSWE (@teortaxesTex (opens in new tab), @teortaxesTex (opens in new tab)).
GPT-6.1 Sol and OpenAI’s DevDay Agent Stack
Independent evals: GPT-6.1 Sol is the new #1 on MathArena (@j_dekoninck (opens in new tab)).
Code Arena: It ranks #3 on WebDev at 1759, 70 points above GPT-6 Sol for the same \$2/\$10 pricing (@arena (opens in new tab)).
Cost per task: Artificial Analysis measures \$0.72 per task at max effort, versus \$3.26 for Astra and \$1.04 for GPT-6 Sol (@ArtificialAnlys (opens in new tab)).
- Source of savings: Sol uses fewer turns and has a lower cache-read price (@ArtificialAnlys (opens in new tab)).
Luna bug fix: OpenAI fixed an image-encoding bug, adding 1 Intelligence Index point to GPT-6 Luna.
Ultrafast inference: OpenAI quotes up to 300 tok/s. SemiAnalysis reports it runs on NVIDIA GPUs at low batch sizes, not on Cerebras (@kimmonismus (opens in new tab)).
- Hands-on report: Generation is about 8x faster, but end-to-end agent tasks speed up only 2–4x because tool latency dominates (@sayashk (opens in new tab)).
Computer use: Gains are largest here, since UI actions respond in milliseconds.
Cost: The tester exhausted a weekly limit in about 2 hours.
- Hands-on report: Generation is about 8x faster, but end-to-end agent tasks speed up only 2–4x because tool latency dominates (@sayashk (opens in new tab)).
Product layer: DevDay introduced dots (persistent agents with their own cloud computers), a Decisions API and computer use (@latentspacepod (opens in new tab)).
Sites: ChatGPT Sites can now host MCP servers and turn them into installable plugins (@mxstbr (opens in new tab)).
Usage limits: Users report one-off credits worth about \$2,500. Others complain that usage limits were cut (@kimmonismus (opens in new tab), @kimmonismus (opens in new tab)).
Other Releases: Embeddings, Image/Video and Open Models
Perplexity contextual embeddings: pplx-embed-v2-context-9b-preview is open on Hugging Face (@perplexity_ai (opens in new tab)).
Method: The model encodes the whole document once and pools chunk vectors afterward. Training distills relevance from a context-compression model instead of using single gold-chunk labels (@denisyarats (opens in new tab)).
Results: It sets a new state of the art on ConTEB. On turbopuffer’s private context-bench it beats voyage-context-4 by 14.4 points in answer recall@10, using 1 KB int8 vectors against 8 KB (@turbopuffer (opens in new tab)).
Cohere Embed 5: The family has Pro and Fast variants in a shared embedding space, so you can index with one and retrieve with the other (@cohere (opens in new tab)).
- Fast tier: Cohere says it beats other fast-tier models by at least 6 points at a third less cost than Pro. Evaluation uses its new RCP-nDCG@10 metric (@cohere (opens in new tab)).
Ideogram 4.5: The editing model targets artifact-free multi-turn edits, with open weights promised (@ideogram_ai (opens in new tab)).
Edit fidelity: Over ten consecutive edits, 94–99% of untouched content stays identical (@fal (opens in new tab)).
Ranking: It is #18 in Image Edit Arena at 1351 (@arena (opens in new tab)).
Video benchmark: Artificial Analysis launched AA-Video-T2V v2.0, judged at 1080p with more than 68K human votes (@ArtificialAnlys (opens in new tab)).
Leaders: Wan 3.0 is #1 at \$12/min. Seedance 2.5 is #2 at \$34.12/min, and MiniMax H3 is statistically tied at \$4.80/min.
Utopai X: This post-train of MiniMax H3 debuts at #2 (@ArtificialAnlys (opens in new tab)).
Open and small models:
Ling-3.1-flash: A 500B model reported close to GPT-5.6 Sol and Opus 5 (@kimmonismus (opens in new tab)). It ranks #2 among open-weight models in Mobile App Arena (@DesignArena (opens in new tab)).
Praxis-1: Runway released an open-weight world-action model and says robotics policy performance scales predictably with third-person video (@agermanidis (opens in new tab)).
Solar Mini 4: Upstage reports 35B total / 3B active parameters. It scores 24 on the Intelligence Index at \$0.10/\$0.40 (@ArtificialAnlys (opens in new tab)).
- Caching penalty: It still costs about 5x Luna per task, because only 48% of its repeated context hits cache versus 99% for Luna (@ArtificialAnlys (opens in new tab)).
Agent Research, Inference and Systems
Context Language Models (Meta): CLMs treat context as an editable file rather than an append-only log, with context-management policies learned in the weights and no external harness (@RulinShao (opens in new tab)).
- Result: They score 65% higher with the same compute on a 24-hour multi-repository agent-swarm task (@arankomatsuzaki (opens in new tab), @natolambert (opens in new tab)).
Adaptive reasoning compute:
TaH2: Lookahead depth supervision teaches the model which hard tokens deserve another loop (@ZhihuFrontier (opens in new tab)).
Gains: It reports +3.4pp accuracy at matched test-time compute and a 53% steeper scaling slope.
Serving: A MiniSGL integration batches requests at different loop depths together.
AutoBenchmark (Meta): The project automates benchmark creation. Human feedback at the ideation stage beats agents working alone, and difficulty transfers to held-out solvers (@jaseweston (opens in new tab)).
Stratego: A Nature paper presents the first superhuman Stratego AI, built on RL and test-time compute under imperfect information (@ssokota (opens in new tab)).
Prefill/decode disaggregation: A steady-state analysis argues that disaggregation raises mean interactivity by about 1/(decode-time fraction) at equal batch size and throughput (@ekzhang1 (opens in new tab), @cHHillee (opens in new tab)).
- Implication: It helps prefill-heavy workloads, not decode-bound low-latency serving.
Compilers and hardware:
DeepSeek on Huawei: DeepSeek released an open-source Ascend toolkit with TileLang optimized for Ascend 950 (@kimmonismus (opens in new tab)).
AI as compiler: A model translates Triton directly to PTX, with a verifier checking correctness, races and deadlocks. Speedups on B200 reach 1.37x on FlashAttention (@Azaliamirh (opens in new tab)).
Vera Rubin: Cognition is the first customer on Vera Rubin via CoreWeave, reporting about 4.8x the token throughput of GB200 at the same decode speed (@cognition (opens in new tab)).
DFlash drafts: New draft models for Ornith-1.5 give up to 2.54x lossless speedups (@ornith_ (opens in new tab)).
Agent sandboxes: Cloudflare rebuilt Containers for agents, with p50 time-to-interactive of 648 ms (6x faster) and snapshots in beta (@mgamache (opens in new tab)).
- AutoRouter: Cloudflare’s model router showed about 30% lower spend in internal tests (@ashleypeacock (opens in new tab)).
Safety, Security and Eval Integrity
Reasoning extraction: OpenAI attributes a core part of a hidden-reasoning extraction campaign to individuals linked to Moonshot AI (@kimmonismus (opens in new tab)).
Scale: OpenAI recorded 16,000 attempts from more than 4,000 users in two days, with related activity across more than 15,000 users.
External researchers: Their attacks kept working on Astra until this week. Patches were hard to propagate across product versions and third-party hosts (@JSchaeff3r (opens in new tab), @jonasgeiping (opens in new tab)).
Criticism: Nathan Lambert argues the vulnerability is the API provider’s responsibility (@natolambert (opens in new tab)).
Distillation defenses: Defenses evaluated without later RL give a false sense of security. RL makes simple attacks effective (@shidan_javaheri (opens in new tab)).
Embedded evaluations: Apollo Research published principles for outside evaluators who receive employee-like access to frontier labs (@ApolloResearch (opens in new tab)).
Cyber evals: On CyberGym-E2E-AA, some frontier models are safety-blocked on more than 85% of tasks (@ArtificialAnlys (opens in new tab)).
- Cost: GPT-6 Luna or MiMo-V2.6-Pro can run about 100 bug hunts in a 1M-line codebase for roughly \$20.
Provenance and transparency:
SynthID Bio: Watermarking for AI-generated proteins is published in Nature, with open-sourced tools (@demishassabis (opens in new tab)).
AI-detector evasion: Opus 5.5 and Astra can rewrite more than 50% of a document without Pangram flagging it (@ValsAI (opens in new tab)).
Agent reports: A new preprint asks how transparent LLM-written reports on agent work actually are (@jennyihuang (opens in new tab)).
Industry and Policy
Factory vs Cognition: Factory removed advisor Chris Degnan, alleging he was confiding in Cognition while attending its board meetings (@matanSF (opens in new tab)).
Hire: Cognition announced Degnan as its CRO the same day (@cognition (opens in new tab)).
Denial: Cognition’s CEO says no Factory information was shared and that Degnan had resigned as an advisor on Monday (@ScottWu46 (opens in new tab)).
Political spending: Greg Brockman dropped a promised second \$25M donation to the Leading the Future super PAC (@teddyschleifer (opens in new tab)).
- Follow-up question: Alex Bores asked whether this also covers anti-regulation groups that don’t disclose donors (@AlexBores (opens in new tab)).
OpenAI finances: NYT reports OpenAI is near \$70B in annualized revenue and in talks to raise \$30B at a \$1.4T valuation, with its IPO pushed to next year (@srimuppidi (opens in new tab)).
Funding: Flow, which builds AI tooling for hardware engineering, raised a \$50M Series B at a \$750M valuation (@parisingh (opens in new tab)).
Top tweets (by engagement)
Gemini 4 Argon introduced; trusted-tester rollout via Fairwind (opens in new tab) — 44.6K
Google: Argon with 1M output limit (opens in new tab) — 36.5K
Factory terminates advisor over Cognition conduct (opens in new tab) — 6.4K
Artificial Analysis: Argon matches Astra at 53 (opens in new tab) — 4.5K
Cognition CEO disputes Factory’s allegations (opens in new tab) — 3.8K
Argon agents freed 300 TiB of memory and drive Rust migrations (opens in new tab) — 3.7K
Ideogram 4.5 for precise multi-turn editing (opens in new tab) — 3.2K
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. GLM-5.3 Cyber Risk and Local Inference Support
GLM-5.3 and the Spread of Advanced Cyber Capabilities \ Anthropic (opens in new tab) (Activity: 785): Anthropic reports that Zhipu/Z.ai’s open-weight GLM-5.3 (opens in new tab) crosses a notable threshold for autonomous cyber capability:
50/410end-to-end V8 exploits on ExploitBench, close to Claude Mythos Preview’s56/410, plus full control-flow hijacks on4%of Anthropic’s internal binary exploitation tasks where prior models were near zero. Anthropic frames the risk as capability + accessibility**: GLM-5.3 is widely downloadable, relatively cheap, and weakly refusal-tuned, with simple jailbreaks reportedly succeeding64–100%of the time and “abliteration” dropping refusals to low single digits with little measured capability degradation.** Top comments were largely hostile to Anthropic’s framing, arguing the post reads as an attempt to suppress a cheaper/open Chinese model near Anthropic’s frontier. One commenter emphasized legitimate defensive use, saying GLM-5.3 is their only practical tool for security testing and improving their own software.Commenters highlight GLM-5.3 as a low-cost, less-restricted model perceived to be close to frontier capability, with one user framing it as useful for “security testing and improvements on my own software” rather than inherently malicious. The technical concern raised is that restrictions by providers like Anthropic could limit defensive cybersecurity workflows that require models willing to analyze potentially sensitive exploit or vulnerability patterns.
One commenter references prior GLM-5.2 models as having helped mitigate a Hugging Face attack, contrasting that with Claude allegedly refusing assistance. The substantive point is that refusal policies may reduce utility in incident response or vulnerability remediation scenarios, while more permissive models can be operationally useful for defensive security tasks.
add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp (opens in new tab) (Activity: 348): Merged
ggml-org/llama.cpp#27773adds GLM-5.3-Flash / GLM5-Next support tollama.cpp, enabling local inference for the 320B hybrid text+vision model. The implementation adds GLM-specific DSA indexing/pooling, hybrid indexed memory, and a newglm5vvision preprocessing/tower path, while reusing Kimi-K3 KDA layers, DeepSeek-style MoE/mHC helpers, MLA-only attention, and DSV4-style SwigLU clamping; validation reports random-model logits matching Transformers across prefill/ubatching/decode and vision embedding agreement around1e-5, with some precision-sensitive tensors left unquantized. Commenters were concerned thatllama.cppmodel support is lagging behind the pace of new experimental architectures, with one noting the effective bottleneck appears to be maintainer availability. A technical compatibility issue was also raised: existing Unsloth quantizations reportedly useglm5nextwhile mainline expectsglm5-next, so current mainline may fail to load those quants.Commenters noted a compatibility issue between the Unsloth quantization PR and the mainline
llama.cppPR: one identifies the architecture/model type asglm5nextwhile the other usesglm5-next, meaning mainlinellama.cppmay fail to load existing Unsloth GLM-5.3-Flash quants without conversion or metadata fixes.There was concern that
llama.cppsupport is lagging behind the pace of new model releases, especially as newer models increasingly use experimental architectures that require bespoke loader/runtime changes before inference and optimization work can land. One commenter framed GLM-5.3-Flash support as taking roughly “another month” after model release, with progress depending heavily on a small number of maintainers.
- Google reports Gemini 4 Argon agents freed more than 300 TiB of data-center memory and are migrating over 800,000 lines of C/C++ kernel code to Rust. In one video-decoder example, agents replaced 32,000 lines of SIMD code with safe Rust; the existing Rust port became 2.7× faster with identical output. Argon access is limited to government users and trusted cyber defenders in the Fairwind Program while Google refines guardrails before broader developer access.
- Context Language Models offer a different agent-context design: treat context as an editable file rather than an append-only log, with context-management policies learned in the model instead of an external harness. The reported result was 65% higher performance at the same compute on a 24-hour multi-repository agent-swarm task.
- OpenAI’s DevDay agent stack includes dots—persistent agents with their own cloud computers—a Decisions API, and computer use; ChatGPT Sites can host MCP servers as installable plugins.
- A speed caveat for agent workflows: an OpenAI hands-on report found roughly 8× faster generation produced only 2–4× faster end-to-end agent tasks because tool latency dominates; gains were largest for computer use, where UI actions respond in milliseconds. Cloudflare’s agent Containers report a 648 ms p50 time-to-interactive (6× faster), with snapshots in beta; its AutoRouter showed about 30% lower spend in internal tests.