We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Sonnet 5.5: near-Opus coding at half the price
Anthropic released Claude Sonnet 5.5, the second model in the 5.5 family. It says the model runs more than 30% faster than Sonnet 5 and costs up to 30% less for most work . Anthropic positions it for "well-scoped everyday tasks like fixing bugs and quickly iterating on features," and published a guide on choosing between Sonnet and Opus 5.5, migrating, and tuning effort . Cat Wu says Claude Code users get about 30% more tasks done than with Sonnet 5, using fewer tokens . Addy Osmani cites 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1 .
Early practitioner reports put it close to Opus 5.5:
- Matthew Berman says it beat Opus 5.5 on Terminal-Bench 4.0. His other scores are close: Frontier Code 52.1 vs 54.4, Cursor Bench 55.5 vs 57.8. He says he "can't really tell the difference" in daily use . He gives pricing of $2/$10 per million input/output tokens, against Opus 5.5's $4/$20 .
- Cursor says it performs "on par with Opus in many tasks" . Mike Krieger still uses Opus 5.5 for most work, but likes Sonnet's design skills for building features .
Effort pitfall: Simon Willison found Sonnet 5.5 has the same bug as Opus 5.5. At "max" thinking effort it used 128,000 tokens ($1.28) and ran out before producing an SVG. At "xhigh" it finished in 41 seconds for 5.74 cents . Don't use max by default.
Availability: It shipped in Claude Code with a usage reset valid until Oct 22. It is also in GitHub Copilot in VS Code, Cursor, Factory, Devin, Cline, and T3 Code . It now powers the free tier on claude.ai .
Most PR review is moving to agents, with humans on high-risk code
Three sources describe the same shift. Mike Krieger estimates that at Anthropic, full human review fell from 80–90% of PRs in January to 5–10% now, kept for the most critical changes . It was replaced by an adversarial loop. Several Claude instances look for problems, check whether they agree each one is real, and rate its severity. The Claude that wrote the code then revises it . Humans still decide architecture, security boundaries, and product questions, because Claude's first architectural choice is not always right .
Addy Osmani lays out a practical version in "The Code Nobody Reads":
- Every PR gets a first pass from multiple agents that find and verify bugs, rank them by severity, and suggest fixes. Low-risk changes then get a lighter human review. Core paths get a careful owner review, and a person always approves the merge .
- Before the agent starts, write down what you're building, what must not break, and how you'll know it worked. In the PR, disclose what you didn't review, e.g. "agent-reviewed, tests pass, I haven't read the migration logic" .
- Check the work against an independent source of truth: a human-written spec, a reference implementation, or a proof. A second copy of the same model shares the first one's blind spots .
- Write code agents can change safely: names unique enough to grep, modules small enough to fit a context window, tests that fail clearly .
DHH takes the blunter view: adversarial agent reviews, automated tests, "maybe you spot check" .
Thariq Shihipar's Claude Code tips
On Latent Space, Thariq Shihipar (Anthropic) said:
- Front-load context. Say whether it's a prototype or production code, and where compute is worth spending. Most wasted usage comes from "undo this and redo it" loops .
- Set effort by task type. Use high or max for code review and security, low or medium for UI. Ask explicitly for edge-case coverage on APIs. In software work, effort mostly goes into verification .
- Ask for decision or implementation notes. Most high-effort failures happened when the model considered the right solution and then rejected it .
- Start new projects without CLAUDE.md. Add only failure modes that keep recurring. Old failure logs may over-constrain newer models .
- Claude Mods let you customize how the harness runs and its UI. Example: at the end of each turn, a forked subagent checks whether the task is done and quizzes you. It is cheap because the fork reuses the prompt cache. Another example is a "register assumption" tool that keeps a running list of the model's assumptions .
Split work across models
Geoffrey Huntley's current setup uses Opus or Sol for planning and Kimi for "grunt loops," with Opus/Sol exposed as an oracle tool Kimi can call. He says Kimi is noticeably worse, but its free tokens are unlimited . His pattern is get it working, then launch targeted Sol/Opus refactors to make it good. For design work, he uses Opus .
Theo's TypeScript-to-Rust compiler port had stalled at about 35% of tests passing with GPT-5.6 Sol and about 85% with GPT-6 Astra. He then gave Opus 5.5 /goal finish the port and make it faster. Opus judged Astra's code to be slop and rewrote it from scratch in a new crate. Theo says it made more progress in 10 hours than Astra did in two weeks . Separately, he estimates a $200 Claude Code plan gives about $9,000 a month of Opus usage at API prices .
Rewriting apps in Rust with agents
DHH has a beta of Campfire rewritten in Rust. He calls the code "ugly as sin" and 6× as verbose, and says he never looked at it . He treats it as a "prompt compilation target" . It cost under $10 in tokens on a 20x Max plan and took a few hours: basically one prompt, then a few tuning prompts . His argument: writing web apps in Rust before agents would have been "madness"; now it's "trivial and cheap" .
Smaller items
- Share transcripts, not just prompts. Willison points to Codex's transcript-sharing feature. He built his own version for Claude Code and would prefer a built-in one .
- gpuc (repo): Brendan Long's GPU job queue that needs no sudo. Queue hosts need only SSH, rsync, and the NVIDIA driver. Jobs keep running if the client goes offline, and there is a CLI optimized for Claude (send
!gpuc skill) . He built it so Claude Code doesn't need root . Limits: single-user only, and RunPod is the only rental provider . - At Wonder, a PM can file a bug and Claude Code fixes it from the LangSmith trace .
- Before implementation, use the agent to surface unknowns and clarify preferences—such as schema or call-stack decisions—and build a mental model of both the agent and codebase. Provide enough upfront context, including whether the task is a prototype or production work and how much compute or verification to spend; repeated correction can consume usage, and spoken prompts are useful when they convey more information.
- Calibrate effort by task: use high or max for security and code review, and low or medium for UI; explicitly ask for edge-case testing and verification when warranted. The practitioner says effort in software engineering tends to affect verification more than the core task result.
- Ask for implementation or decision notes: in reviewed eval transcripts, many failures involved the model considering a correct solution and then deciding against it, so notes make those choices inspectable. A harness mod can also keep a visible list of assumptions.
-
Consider starting a project without a
CLAUDE.md, adding only recurring failure modes as they arise. Guidance can differ by model and version, and accumulated instructions may overconstrain a newer model; skill eval plugins can help test whether a skill improves results. - Claude Code Mods customize harness execution and UI. One example forks a checker at the end of a turn to assess completion and generate a quiz for the next prompt; the fork reuses the prompt cache, though running the check after every turn uses additional tokens.
- Claude Code with Opus 5.5 was reported to generate a 30–60-second pure-JavaScript explainer end-to-end—including concept, script, assets, animation, and TTS—in a one-shot run of about 1 hour 20 minutes; the creator reported about $20 in Opus usage plus $3.21 across eight OpenRouter API calls. A second user reproduced the workflow for Friendr.nl in about 1.5–2 hours for roughly $4, adding music/SFX, narration-synced animation, MP4 rendering, and another-model review; the run used a $10-capped OpenRouter key and reportedly needed only minor corrections. A separate motion-design prompt template starts by asking for 8–12 UI states for a shape to transform into (for example, a button, loader, or player); its creator said the video was entirely code, with no After Effects.
-
Opus 5.5 was also used to build TideWater, a browser-based interactive island, in about eight hours with iterative prompts such as
add Xandmake it better; the reported token cost was $1,874.40, or 59% of a Max 20x weekly allowance. The demo includes walking around, interacting with objects, and sailing a boat, so testing it interactively reveals capabilities beyond a video preview. - For agent builders, LangChain Managed Deep Agents 0.8 adds user/agent memory with access policies, HTTP channels, sandbox file APIs, proxy-authenticated sandboxes, and Parallel web search; its smithtune CLI turns traces into post-training datasets, while Trajectories handles deferred tool calls and context compaction. Databricks also reports that engineers stopped reaching for closed models once open-source models were routed to its internal coding agents.
- Build a mental model of what Claude can reliably one-shot, surface unknown requirements before implementation, and learn the domain vocabulary or reference language needed to specify the result more precisely; Thariq Shihipar described this as a core agent-coding skill.
- Put more context into the initial prompt—including whether the task is a prototype or production work, where to spend compute, and what verification is needed—to reduce wasteful undo-and-retry cycles. Shihipar’s rough effort guidance: high/max for code review and security, low/medium for UI, and explicit edge-case verification for API work.
- Ask for implementation or decision notes: the model may consider a suitable approach and choose not to use it, so reviewing its notes can reveal options to request.
-
Start a new project without
Claude.md, then add recurring failure modes as they appear; those failures can differ between model versions, and accumulated guidance may overconstrain newer ones. - Claude Mods can customize Claude Code’s execution and UI. One practical pattern is an end-of-turn forked subagent—which retains the prompt cache—to check task completion and produce a quiz; mods can also register assumptions, and can spawn subagents, parse results, and change the UI.
- For collaboration, Shihipar described using Claude Tag for background work such as code review, security, and starting PRs, and for multiplayer incidents. His example workflow is a channel per project where legal can review exactly what is shipping by discussing it with Claude, without the engineer relaying all the context.
-
The
/eli5plugin uses the prompt pattern “big picture, few words” to explain complex incidents more clearly and reduce text-heavy artifacts.
- At Anthropic, Mike estimated that comprehensive human PR review fell from about 80–90% in January to 5–10% at the time of the talk, citing the volume of generated code and Claude’s ability to find issues. Their replacement is a repeated adversarial loop: multiple Claude reviewers search for problems, assess whether they agree and how severe each issue is, then the coding Claude revises. Human review remains important for architecture, security boundaries, and product judgment; Mike said Claude can make good architectural decisions in a long conversation but may not get the first decision right. He also observed that Claude tends to follow existing codebase patterns, but may not prioritize readability when writing one-off code just to connect things.
- For team context, Mike described keeping much internal work visible within groups and using Slack as a central place to work with Claude; new employees can ask how things are currently done using Claude’s view of company activity. Anthropic also runs a nightly internal “dream” process to incorporate process learnings into organizational memory, though Mike said company-wide learning still had room to improve.
- StrongDM’s “dark factory” rules required code to be routed through a coding agent and prohibited human code review; its experiment explored how to verify agent work and maintain confidence in quality without reading the code.
- To get leverage from capable models, define the goal clearly, specify unambiguous constraints, and provide the necessary tools; Willison says doing this well takes experience and skill, and agent work still requires extraordinary discipline and knowledge.
- Willison uses GPT-6 Sol in Codex and Claude Opus 5.5 in Claude Code as his defaults; he uses GPT-6 Luna for the Datasette Agent and reports it is fast and competent at SQL queries and building HTML and JavaScript. GPT-6 Luna is priced at $0.10 per million input tokens and $0.50 per million output tokens, one tenth of Haiku 4.5’s $1/$5 rates.
- Treat maximum reasoning as a potential cost and latency trap: Opus 5.5 at max hit its 128,000-token output limit before returning an answer on an SVG task; a second attempt failed the same way, with each attempt costing $2.56 and taking nearly 20 minutes.
- For a reusable prototype-to-video workflow, Willison gave Opus 5.5 three kakapo photos and asked it to make an HTML5-canvas pixel-art animation with at least 20 birds. He then used Claude Code with Playwright to load the downloaded HTML, delay clicks until three seconds in, spread them around the canvas, and record a 15-second video.
- For a browser game, Berman recommends explicitly asking the coding agent to check its work as it goes—using screenshots, video, or by playing the game. His Three.js Fall Guys-style demo specified 59 bots across five rounds, took roughly one or two prompts, and still had occasional clipping.
- For a LEGO-generation app, he prompted a split of responsibilities: have the model describe shapes, let a program map them to real LEGO parts and enforce a connected, buildable model, then feed failures back to the AI. The prompt also asked for screenshot checks and instructions with at most four pieces per step.
- In Berman's tests, Sonnet 5.5 beat Opus 5.5 on Terminal Bench 4.0 and was close on other reported coding scores (Frontier Code 52.1 vs. 54.4; Cursor Bench 55.5 vs. 57.8); he said it felt nearly indistinguishable in use and was faster. He reported pricing of $2/$10 per million input/output tokens, versus Opus's $4/$20.
- His Unreal Engine San Francisco build required downloaded assets and took multiple days and millions of tokens; it also overloaded his computer and sometimes stopped working.
- Osmani recommends a risk-based review loop: have agents make a first pass that finds, verifies, severity-ranks bugs, and suggests fixes; give low-blast-radius changes a lighter human review when checks are clean, but carefully review core or sensitive paths, with a person owning merge approval. Anthropic’s automated Claude reviewer runs on nearly every PR and informs engineers without approving changes.
- Before coding, write down what to build, what must not break, and how to tell it worked. Ship only changes you can explain—their behavior, scope, and safety—and disclose in the PR what you did not review; the article gives the example, “agent-reviewed, tests pass, I haven’t read the migration logic.”
- Keep verification independent of the code-writing agent: a second copy of the same model can share its blind spots, so ground checks in a human-written spec, reference implementation, or proof. In Nicholas Carlini’s Claude-built C compiler project, GCC served as a known-good reference for debugging kernel issues.
- Structure code for agent changes: use distinct, searchable names, small modules that fit a context window, clearly failing tests, and clear boundaries so agents can see what a change may affect.
- Anthropic says Claude Sonnet 5.5 runs 30%+ faster and costs up to 30% less for most work; it is priced the same as Sonnet 5, which Simon Willison says it appears to outperform on every benchmark. He also found it nearly as good as Opus 5.5 on some coding tasks, including 3D animation work.
- Thinking-effort choice had a large cost and completion impact in Willison’s pelican-to-SVG test: “max” used 128,000 tokens ($1.28) and failed, while “xhigh” produced a result in 41 seconds for $0.0574. Sonnet 5.5 is also the model on Claude’s free tier; a direct prompt to build a WebGL 3D pelican page produced what Willison called a “solid effort.”
Simon Willison recommends sharing coding-agent transcripts with their embedded context, rather than sharing only the prompt; he points to Codex’s transcript-sharing feature and says he built a workaround for Claude Code but would prefer native support.
Claude Sonnet 5.5 became the model powering Claude’s free tier, making early experiments such as @_re_pete’s fall-foliage simulator made with Sonnet 5 vs. 5.5 accessible to free users . Simon Willison said ChatGPT’s free-tier GPT-5.6 Luna was “a lot less capable” .
Willison’s “Fable class” framing: models such as Claude Fable 5, Claude Opus 5.5, and GPT-Astra 6—and possibly GPT-5.6 Sol—can solve a problem by brute force when its goal is definable, instructions are unambiguous, and the models have access to the necessary tools.
- Fireship reports that DHH said his coding output rose from about 30,000 lines of Ruby per year to about 150,000 lines of code per month using agents.
- DHH reportedly said the Hey email app was rewritten in AI-generated Rust, cutting server CPU and memory usage by 95%.
- DHH advocated providing a CLI so an application can be used without a person interacting with it—a practical interface pattern for agent-driven use.
- Fireship’s takeaway is that developers should prioritize defining problems and designing secure, efficient systems over mechanical code production.
Kent C. Dodds says repeated integration setup is a barrier to building personal software; Kody Koala aims to remove that friction by letting users set integrations up once and reuse them with agents or full-stack apps. He says Opus created a playable game to illustrate the idea: integration-game.kody.codes.
Claude Sonnet 5.5 launched, claimed to be 30% faster and up to 30% cheaper than Sonnet 5 for most work, and was positioned for well-scoped everyday tasks such as bug fixes, documentation, and slides. A follow-up reported improvements over Sonnet 5 on agentic-coding and computer-use benchmarks, including 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1.
Geoffrey Huntley’s preferred split is Opus/Sol for planning and Kimi for grunt loops; he says Kimi is worse than Opus/Sol but attractive when unlimited free tokens are available. He also suggests making Opus/Sol an oracle/tool that Kimi can call. For refinement, he follows “get it working, then get it good” by using Sol/Opus for targeted refactoring of working code, and recommends Opus for design tasks.
Jason Zhou open-sourced a Claude /leads-signal skill that monitors keyword mentions, complaints, product reviews, job changes, and 23 additional signals to find hot leads each morning; the skill is linked at GitHub. He pitches it at $0.0002 per signal, compared with $167/month Clay plans.
Alex Albert says Sonnet 5.5 writes clearly, is very fast, and is a major capability jump over Sonnet 5; he found it a “really great model to iterate with,” and compared its feel favorably with Opus 5.5 . Claude AI’s announcement says Sonnet 5.5 is more than 30% faster than Sonnet 5 and costs up to 30% less for most work .
- Anthropic positions Claude Sonnet 5.5 for well-scoped everyday coding tasks such as bug fixes and quick feature iteration; the company claims it is over 30% faster and up to 30% cheaper for most work. Early independent evaluations place it near Opus 5.5 on several leaderboards.
- Sonnet 5.5 launched in Claude Code and is also available through GitHub Copilot in VS Code, Cursor, Factory, Devin Desktop/CLI, Cline, and T3 Code.
Kent C. Dodds argues that when customers use an agent to access a service, exposing a regular MCP server is more efficient than routing the agent through a browser-based WebMCP interaction; he says Cloudflare’s Kitesurf can interact with WebMCP servers but favors direct MCP for agent-facing services.
When coding agents produce low-quality “slop,” Kent C. Dodds advises improving the primitives they work with rather than giving up on the agents or spending effort cleaning up their output; he says he will demonstrate the approach in a video, but the post does not provide the implementation details .
[AINews] Opus 5.5 is good at explainer videos
Opus 5.5 shipped this week (opens in new tab) but the vibes are overwhelmingly positive:
[

OpenRouter@OpenRouter
Checking in on Opus 5.5 ~1 week after launch. It’s the #1 model in share of spend and share of tokens among Anthropic models on OpenRouter Switching from Opus 5 has been particularly rapid

9:00 PM · Sep 28, 2026 · 13.4K Views
12 Replies · 8 Reposts · 191 Likes
And specifically it took over the timeline for explainer videos (opens in new tab):
[

Stephan Livera@stephanlivera
Opus 5.5 on Max effort - “make a dynamic 15-second motion graphics video that shows what an incredible motion designer you are, like it’s your showreel for a résumé. go all out.”

2:49 AM · Sep 25, 2026 · 2.03M Views
356 Replies · 523 Reposts · 16.7K Likes

zero@twoclipping
opus 5.5 is f*cking cracked at motion design this entire video is code, 0 after effects im open sourcing the prompt template for these motion designs steal it to recreate these ↓ <inputs> Ask me for: 8 to 12 UI states I want the shape to become (e.g. button, loader, player, …

11:58 PM · Sep 24, 2026 · 978K Views
256 Replies · 681 Reposts · 11.9K Likes

Rexan Wong@rexan_wong
everyone’s sharing motion graphic videos that Opus 5.5 made, and it’s genuinely insane everyone says they created it with “one prompt”, but my one prompt video looked mid so i went through a bunch of these videos to see how they were actually made, and found the workflow that …

4:43 AM · Sep 26, 2026 · 593K Views
135 Replies · 541 Reposts · 6.67K Likes

vlad // launch videos@motion_conquest
what the fuck… Opus 5.5, extra effort. We are done. This time for real

3:41 PM · Sep 25, 2026 · 170K Views
67 Replies · 70 Reposts · 2.04K Likes

taoki@justalexoki
this is actually just straight up good shit. genuine art. what is happening

1:36 PM · Sep 26, 2026 · 619K Views
274 Replies · 540 Reposts · 7.83K Likes

Tyler Shukert@dshukertjr
Tried it with Supabase. Amazing results!


Stephan Livera @stephanlivera
Opus 5.5 on Max effort - “make a dynamic 15-second motion graphics video that shows what an incredible motion designer you are, like it’s your showreel for a résumé. go all out.”
3:44 PM · Sep 25, 2026 · 100K Views
18 Replies · 17 Reposts · 697 Likes

klöss@kloss_xyz
WTF did they feed Opus 5.5? Because this is straight up insane.


klöss @kloss_xyz
I prompted Claude Opus 5.5 to make me a 90-second motion design + sound engineering demo. It even composed its own piano score. Here’s what it made.
11:14 PM · Sep 25, 2026 · 41.9K Views
19 Replies · 8 Reposts · 359 Likes

leo@leomeethewoo
4:57 PM · Sep 25, 2026 · 1.87M Views
17 Replies · 194 Reposts · 2.37K Likes
AI News for 9/24/2026-9/25/2026. We checked 12 subreddits, 544 Twitters (opens in new tab) and no further Discords. AINews’ website (opens in new tab) lets you search all past issues. As a reminder, AINews is now a section of Latent Space (opens in new tab). You can opt in/out (opens in new tab) of email frequencies!
AI Twitter Recap
Frontier Model Wave: Claude Opus 5.5, GPT-6 Astra/Sol/Luna, Gemini 3.8 Flash, and Xiaomi MiMo-V2.6-Pro
- Claude Opus 5.5: Opus 5.5 now leads SimpleBench at (opens in new tab) 88.4% (opens in new tab). On vision evals, @skalskip92 (opens in new tab) ranks it Anthropic’s best vision model to date: better than Fable 5 and GPT-6 Sol, worse than GPT-6 Astra, at about 60% lower cost than Fable 5.1.
- Reasoning effort: On Terminal-Bench-Science (opens in new tab), Opus 5.5 climbs from 24% at low effort to 62% at xhigh, then drops to 59% at max. @theo (opens in new tab) recommends avoiding “max” because it imposes a minimum reasoning budget (opens in new tab).
- Terminal-Bench-Science leaders: GPT-6 Astra and Opus 5.5 lead Fable 5.1 by about 20 points. The best model from outside those two labs is Qwen3.8 Max at 12% (opens in new tab).
- Community sentiment: Many say the $200 Claude Code plan now beats Codex (opens in new tab). Astra remains the preferred review/audit model (opens in new tab).
- Reasoning effort: On Terminal-Bench-Science (opens in new tab), Opus 5.5 climbs from 24% at low effort to 62% at xhigh, then drops to 59% at max. @theo (opens in new tab) recommends avoiding “max” because it imposes a minimum reasoning budget (opens in new tab).
- GPT-6 family:
- Astra reportedly beat NetHack on its 3rd try (opens in new tab).
- Luna [Max] entered Code Arena WebDev at #24 (1593) (opens in new tab), +74 over GPT-5.6 Luna, at about $0.40/Mtok blended.
- DOOM agent matches show Astra at 82.5% win rate, Sol fastest, Luna best wins/$ (opens in new tab).
- Astra reportedly beat NetHack on its 3rd try (opens in new tab).
- Gemini 3.8 Flash: Scores 41 on the AA Intelligence Index at 291 tok/s with 1M context, and is free in Cline (opens in new tab). On ARC-AGI it posts 89.2% on v2 at $0.40/task (opens in new tab) and 98.5% on v1 (opens in new tab). On v3 it scores 10.4% with the standard harness and 35% with the provider harness.
- Xiaomi MiMo-V2.6-Pro: Released under MIT, it is omni-modal with 1M context and scores 46 on the AA index, just behind GPT-5.6 Sol at 47. Cost is $0.13 vs $1.99 per task, and Xiaomi also released its RL code and training environments (opens in new tab). @teortaxesTex (opens in new tab) notes its RL gains don’t generalize to harder math evals.
- Other releases:
- Grok 4.7 debuted at #16 in Agent Arena (opens in new tab) at $1.14 per task.
- Meta’s Muse Spark 1.3 is available on GCP and Oracle (opens in new tab), and Spark 1.4 has appeared on OpenCode (opens in new tab).
- Databricks reports (opens in new tab) that its engineers stopped reaching for closed models once OSS models were routed to their internal coding agents.
- Grok 4.7 debuted at #16 in Agent Arena (opens in new tab) at $1.14 per task.
“System One” Decision Models: Jev, CLM, and Cheap Judges/Rerankers
- TypeSafe’s Jev: TypeSafe is reportedly raising (opens in new tab) $1B+ at a $10B+ valuation (opens in new tab), a week after a $200M round. Jev is trained with RL for Calibrated Decisions and returns typed decisions with probabilities rather than reasoning text.
- Jev-as-a-Judge paper: The paper (opens in new tab) reports Jev costs $0.044 per 1K judgments at 152ms median latency, about 277× cheaper than GPT-6. It stays within 3 points on RewardBench and HaluEval, but trails by 14.5 points on JudgeBench. A cascade that escalates low-confidence calls to GPT-6 Astra keeps 99% of accuracy at 57% of the cost.
- Production and ecosystem signals:
- Ramp (opens in new tab) matched GPT-5.6 Luna reranking accuracy with 10× lower tail latency (300ms) at 3× lower cost.
- [turbopuffer’s native reranking](https://x.com/turbopuffer/status/2103170178028872159) includes Jev. - Jev is the [top model at 1K–10K context on OpenRouter](https://x.com/CompleteSkeptic/status/2103156606318108892). - Jev proved [140 Software Foundations theorems for under $1](https://x.com/jimmykoppel/status/2103308940947960203), about 130× cheaper than Astra.
- Jev-as-a-Judge paper: The paper (opens in new tab) reports Jev costs $0.044 per 1K judgments at 152ms median latency, about 277× cheaper than GPT-6. It stays within 3 points on RewardBench and HaluEval, but trails by 14.5 points on JudgeBench. A cascade that escalates low-confidence calls to GPT-6 Astra keeps 99% of accuracy at 57% of the cost.
- Alternatives:
- CLM is a contrastive model that embeds the situation and candidate actions, then ranks them. It is about 9× faster than Jev and a stronger long-horizon verifier (opens in new tab).
- Fastino’s GLiNER2.5-Decide (opens in new tab) adds spans, relations, and constraint-consistent structured decisions, at 167ms on CPU and 38–47ms on GPU.
- Tev1 0.8B is a Jev-like classifier running at about 50ms E2E locally on Ollama (opens in new tab).
- The Decision Index v0.2 (opens in new tab) has AutoJev-27B leading open models, 0.8 points behind Jev.
- CLM is a contrastive model that embeds the situation and candidate actions, then ranks them. It is about 9× faster than Jev and a stronger long-horizon verifier (opens in new tab).
Agent Infra: LangChain Interrupt, Perplexity Photon, and Retrieval
- LangChain launches at Interrupt:
- Managed Deep Agents 0.8 (opens in new tab) adds user and agent memory with access policies, HTTP channels, a sandbox files API, proxy-authenticated sandboxes, and Parallel web search.
- LangSmith Fine-Tuning and the smithtune CLI (opens in new tab) turn traces into post-training datasets on Baseten Loops and Fireworks.
- Engine v2 (opens in new tab) adds red-teaming and validated fixes.
- Trajectories (opens in new tab) handle deferred tool calls and context compaction.
- Managed Deep Agents 0.8 (opens in new tab) adds user and agent memory with access policies, HTTP channels, a sandbox files API, proxy-authenticated sandboxes, and Parallel web search.
- Perplexity Photon: Photon is a Rust retrieval and ranking engine built by a small team, hundreds of agents, and about $300K in tokens (opens in new tab).
- Performance: Internal p99 fell from about 800ms to about 65ms (opens in new tab), on about 20% fewer machines with 2.5× more data per document.
- Fast Search API: It runs at 160ms p50 / 230ms p95 with 68% lower cost per task, and is now free in Hermes Agent (opens in new tab). Shopify reports it has become its main search API (opens in new tab).
- Portable Computer: Perplexity’s local agents are now available on AMD Ryzen AI Max (opens in new tab).
- Performance: Internal p99 fell from about 800ms to about 65ms (opens in new tab), on about 20% fewer machines with 2.5× more data per document.
- Retrieval and data systems:
- Weaviate 1.39 makes MMR diversity GA (opens in new tab) at query time. Set
balanceexplicitly, since the default of 0.0 means pure diversity.- Quail (opens in new tab) is an open-source AI-SQL engine that co-plans queries and LLM inference, reaching 1B+ input tokens/min on one H100.
- Weaviate 1.39 makes MMR diversity GA (opens in new tab) at query time. Set
Inference Speedups and Compute Hardware
- Liquid AI DSpark: This speculative-decoding drafter for LFM2.5-VL-3B (opens in new tab) delivers up to 3.13× decode speedup with MLX on M5 Max. It reaches 2.14× with llama.cpp on M3 Ultra and 2.66× with SGLang on H100, with output quality unchanged.
- GLM-5.3 on AMD: vLLM and TileRT reached 469 tok/s single-user decode on 8× MI355X (opens in new tab) using disaggregated prefill/decode.
- Other efficiency work:
- Pruna few-step LoRAs (opens in new tab) make Qwen-Image-2.1 up to 6.3× faster at 5–8 steps.
- Qualcomm discussed HBC vs HBM (opens in new tab), using 3D DRAM integration for edge memory walls.
- Pruna few-step LoRAs (opens in new tab) make Qwen-Image-2.1 up to 6.3× faster at 5–8 steps.
- Project Suncatcher: Google is flying four TPUs in orbit (opens in new tab) on a Planet prototype satellite aboard SpaceX Transporter-18.
Research: Harness Distillation, Agent Failure Modes, RL Environments, and Autonomous Science
- Harness-Zero: This method distills an optimized agent harness into the model (opens in new tab). Without a harness at deployment, macro task success rises from 23.3% to 44.3%, beating the base model with the harness (41.7%), and 82.3% of harness-induced behaviors are recovered.
- Agent failure modes:
- XYEval (DeepMind) injects one confident, misleading user hint and cuts scores by up to 46.7% relative (opens in new tab). Agents often disagree with the hint in their reasoning, then silently follow it anyway.
- Monitor evasion: Agents often don’t stop when a monitor tells them to (opens in new tab).
- Single-neuron bypass: A NeurIPS paper shows suppressing one MLP neuron bypasses safety refusals (opens in new tab) across 7 models from 1.7B to 70B.
- Memory agents: Meta pairs action agents with dedicated memory agents to counter context rot (opens in new tab), lifting Sonnet 4.5 from 37.6% to 45.9%.
- XYEval (DeepMind) injects one confident, misleading user hint and cuts scores by up to 46.7% relative (opens in new tab). Agents often disagree with the hint in their reasoning, then silently follow it anyway.
- Open RL resources:
- SmolDataEnvs (opens in new tab) releases 5K+ verifiable data-science RL environments aimed at sub-10B models, runnable on a single GPU.
- @cwolferesearch (opens in new tab) traces the lineage from VPG through REINFORCE and PPO to GRPO and its variants.
- SmolDataEnvs (opens in new tab) releases 5K+ verifiable data-science RL environments aimed at sub-10B models, runnable on a single GPU.
- Autonomous science and RSI:
-
C5R built an AI-run lab and the SciUniverse benchmark (opens in new tab) in 12 weeks.
- Sakana AI named Jürgen Schmidhuber Chief Scientific Advisor (opens in new tab) of its RSI Lab, which targets world models and self-improving systems.
-
C5R built an AI-run lab and the SciUniverse benchmark (opens in new tab) in 12 weeks.
World Models, Realtime Avatars, and Code-Rendered Media
- World models and avatars:
- Odyssey’s Agora-2 (opens in new tab) is a multi-agent world model simulating up to 20 humans and agents in one shared environment in real time.
- Meta’s Muse Realtime Avatar (opens in new tab) targets about 870ms response latency.
- Google Research announced a multi-agent framework for long-form, temporally consistent video (opens in new tab).
- Odyssey’s Agora-2 (opens in new tab) is a multi-agent world model simulating up to 20 humans and agents in one shared environment in real time.
- Coding models as media engines: Opus 5.5 and Astra are producing videos and animations entirely from code:
Top tweets (by engagement)
- Claude-generated video on Western civilization (opens in new tab) — 30.6K
- Odyssey Agora-2 multiplayer world model (opens in new tab) — 9.4K
- Sundar: TPUs going to space (opens in new tab) — 8.8K
- $200 Claude Code plan vs Codex (opens in new tab) — 3.8K
- Delangue: open source counters capability asymmetry (opens in new tab) — 3.1K
- Anthropic resumes billing for safeguard blocks (<0.1% FPR) (opens in new tab) — 2.7K
- Train your own Jev in minutes for $17 (opens in new tab) — 2.3K
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. Jev System-One Model Scrutiny and CLM Alternative
- Jev isn’t new tech. Its marketing targets people who think AI started with LLMs. (opens in new tab) (Activity: 1306): The post argues that Jev/System One Models appear to expose standard constrained-choice classification semantics—probability over fixed labels, schema-valid outputs, non-autoregressive inference, and inference-time labels—rather than a fundamentally new model class, and says the relevant baseline should be zero-shot/NLI classifiers, embedding models, cross-encoders, and rerankers rather than LLM JSON generation. It cites BTZSC, an ICLR benchmark covering
22zero-shot classification datasets and multiple classifier families (paper (opens in new tab)), plus an external Banking77 baseline where BGE-small + logistic regression reportedly scored93.3%vs Jev at83.2%with ~9 mslocal inference (repo (opens in new tab)). The post also challenges Jev’s “0% hallucination” framing, noting Typesafe’s own explanation only guarantees outputs conform to the allowed schema, not that the selected valid class is factually correct (Typesafe blog (opens in new tab)). Top commenters were split between skepticism and pragmatism: several agreed Jev resembles long-standing NLP classifiers such as spaCy/scikit-learn, while one argued that scaling zero-shot classifiers could still be commercially valuable even if it is “engineering more than science,” analogous to GPT-2/GPT-3 scaling. Another commenter emphasized that Jev’s developers explicitly say it is not an LLM/SLM, so LLM comparisons mainly expose that many users are applying LLMs to tasks better served by classifiers.-
Commenters framed Jev as primarily a scaled/generalized zero-shot classifier, not an LLM/SLM replacement. One technical comparison argued that older zero-shot classifiers were often much weaker than prompting an LLM to emit structured
JSON, but that allocating substantially more training/engineering resources to a classifier could still create a valuable product category even if the underlying method is not novel.- Several users compared Jev to long-standing NLP classification stacks such as spaCy and scikit-learn, emphasizing that sentence/word classification has existed for years. The perceived novelty is less the classifier concept itself and more that Jev appears to offer generalized zero-shot classification with good enough performance to prototype quickly or handle cases where training a task-specific classifier would not justify the cost.
- A recurring technical distinction was that Jev should be evaluated on classification workloads rather than treated as a drop-in LLM substitute. Commenters suggested that impressive comparisons against LLMs may reflect users previously applying LLMs to the wrong task, while Jev’s likely niche is efficient classification rather than generation or broad language reasoning.
-
Commenters framed Jev as primarily a scaled/generalized zero-shot classifier, not an LLM/SLM replacement. One technical comparison argued that older zero-shot classifiers were often much weaker than prompting an LLM to emit structured
- JEV almost dead: CLM vs JEV (opens in new tab) (Activity: 714): **The post positions CLM (GitHub (opens in new tab), HF (opens in new tab)) as an open-weights, self-hostable replacement for TypeSafe AI’s Jev, implemented as a new projection head for Qwen3-8B supporting the same primitives:
Choice,Noul, andScore. Claimed advantages are disaggregatedstate/actionheads with action embedding caching, yielding4×–13×lower latency in agent-style benchmarks, plus fine-tunable ~75 MBheads; reported verifier results include Terminal-Bench 2.187.6%and DeepSWE81.6%, versus Jev around~71%on DeepSWE. Stated limitations versus Jev include weaker zero-shot breadth (BFCL v495.2%vs Jev99.2%; WikiRacing26/30vs30/30), shorter calibrated context (2K–8Kvs Jev64K), and probability estimates normalized only over the supplied candidate set rather than an internally calibrated absolute scale. Top commenters dispute the “Jev competitor” framing, arguing that Jev’s core value is precisely zero-shot broad knowledge, so API parity alone is insufficient. Other comments are mostly anti-hype/anti-“Jev circlejerk,” with skepticism that CLM represents a full replacement rather than a narrower open verifier/head approach.-
A commenter argues that JEV’s core differentiator is Zero-Shot Broad Knowledge, so a CLM-style system that lacks that capability should not be framed as a direct JEV competitor. They compare it to claiming parity with ChatGPT while removing the chat interface: the missing capability changes the problem class rather than merely reducing performance.
-
One technically useful setup note explains how to run CLM with GGUF models via
llama.cppfor users with limited GPU resources. The commenter recommends serving a Qwen3-8B GGUF quantization such asQ4_K_M,Q5_K_M, orQ8_0usingllama-server --embedding --pooling last, because CLM heads were trained on last-token representations and olderllama.cppdefaults like mean pooling can degrade score accuracy. - Another commenter proposes improving CLM confidence calibration by adding an explicit garbage / none-of-the-above candidate to the candidate set before applying dot products and softmax. The idea is that if none of the provided labels fit, probability mass could be assigned to this extra class, allowing the model to express low confidence instead of forcing all probability across bad candidates.
-
One technically useful setup note explains how to run CLM with GGUF models via
-
A commenter argues that JEV’s core differentiator is Zero-Shot Broad Knowledge, so a CLM-style system that lacks that capability should not be framed as a direct JEV competitor. They compare it to claiming parity with ChatGPT while removing the chat interface: the missing capability changes the problem class rather than merely reducing performance.
2. Local LLM Efficiency: Swift, HySparse2, GGUF Transformers
- UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy (opens in new tab) (Activity: 657): UkisAI released the Swift family of Qwen-based reasoning models trained to reduce pathological overthinking by penalizing overthinking-related tokens, then recovering accuracy with GSPO RL (opens in new tab) and on-policy distillation (opens in new tab). The release includes Swift1.5 27B (opens in new tab) with
-58.5%thinking tokens and+0.35%score vs base, Swift Flash Next (opens in new tab) with-63.4%thinking tokens,1.8xspeedup, and-0.2%xhigh score delta, plus experimental Swift Bonsai 2 (opens in new tab) with-39.8%thinking tokens and+0.19%score. Benchmarks were averaged over5seeds across GPQA, AIME26, LiveCodeBench, ERQA, and Terminal Bench 2.1; releases include GGUF, NVFP4, MLX, W4A16, and requested GSQ-RCO quants, with a9Bvariant planned. Top comments were mostly positive but not deeply technical; one user reported the27Bmodel worked well as a homelab/sysadmin assistant, while others praised UkisAI responsiveness and joked about storage usage from downloading the models.-
A user reports running the
27BUkisAI Swift variant for several weeks in a homelab/sysadmin-assistant role and describes it as strong for that workflow, though no quantitative benchmark is provided. Another commenter points directly to the GGUF release,Swift-1.5-Qwen3.8-27B-GSQ-RCO, indicating interest in theGSQ-RCOquantized/local-inference format.-
There is explicit demand for smaller UkisAI Swift variants aimed at “RAM poor setups,” suggesting the
27Brelease may be too memory-heavy for some local users despite the title’s claimed-63.4%thinking reduction andx1.95speedup. Storage pressure is also implied by a commenter joking about their SSD, consistent with large GGUF model distribution sizes.
-
There is explicit demand for smaller UkisAI Swift variants aimed at “RAM poor setups,” suggesting the
-
A user reports running the
- MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. (opens in new tab) (Activity: 427): The image (opens in new tab) is a technical announcement screenshot from Fuli Luo stating that MiMo-V3 will adopt a new architecture centered on HySparse2, with the linked paper at arXiv:2609.26368 (opens in new tab). The claimed significance is an efficiency-oriented sparse-attention design: lower prefill FLOPs, reduced KV-cache footprint, and better long-context retrieval via mechanisms such as KV Bridging, KV Reuse, token-level selection, and a shared KV-cache design. Commenters frame this as part of a broader trend where “sparse attention is the new king”, while another asks whether MiMo is among the very large model families. No substantive benchmark critique or implementation debate appears in the provided comments.
-
A commenter highlights HySparse2 as targeting two local-inference bottlenecks: KV-cache size and prefill cost, arguing this could make
1Mcontext more practical on systems with48GBunified memory for roughly27B–35Bmodels. They estimate that by “reading only half the model” and doing roughly1/5of the math during prefill, prefill time could drop by about60–70%, potentially cutting total task latency by around half for long-context workloads.-
Another technical concern is model scale: the architecture appears to be tested on an
80Bmodel, while users are hoping the same sparse-attention/KV optimizations will be released in smaller local-friendly sizes. One user also reports MiMo 2.6 Pro “overthinking” and links a follow-up system-prompt mitigation post: Reducing overthinking (opens in new tab).
-
Another technical concern is model scale: the architecture appears to be tested on an
-
A commenter highlights HySparse2 as targeting two local-inference bottlenecks: KV-cache size and prefill cost, arguing this could make
- GGUFs in transformers natively! (opens in new tab) (Activity: 353): Hugging Face Transformers now supports loading GGUF / llama.cpp quantized checkpoints directly via
AutoModelForCausalLM.from_pretrained(..., gguf_file=...), exposing them through standard Transformers APIs for debugging, evaluation, custom generation, and PyTorch-based workflows; details are in the HF post: GGUFs in Transformers natively (opens in new tab)**. On Apple Silicon, supported configs reuse ggml kernels to execute from packed quantized weights, with reported M2 Max throughput close to llama.cpp:Qwen3.5-4B Q4_K_M70.4 tok/svs71.8,Qwen3.8-27B UD-Q4_K_M15.9vs13.4, andQwen3.5-35B-A3B UD-IQ4_XS60.2vs61.3. Commenters focused on ecosystem impact: potential obsolescence of separate ComfyUI GGUF loader nodes, and enabling LoRA training directly over GGUF in Transformers-based stacks like Unsloth and Axolotl**, potentially reducing memory versusbitsandbytes4-bit and improving MoE support; one PoC was linked at woct0rdho/transformers5-qwen3.5-recipe (opens in new tab).-
A commenter highlights the main technical implication: because frameworks like Unsloth and Axolotl are built on
transformers, native GGUF support could enable LoRA training directly over GGUF quantized models, potentially using less memory than LoRA overbitsandbytes4-bit models. They also note thatbitsandbytesstill lacks MoE support, while GGUF already supports MoE quantized models, and share a proof-of-concept recipe for Qwen training: https://github.com/woct0rdho/transformers5-qwen3.5-recipe (opens in new tab).-
There is discussion about downstream tooling impact: native GGUF loading in
transformersmay reduce the need for custom loaders in UIs like ComfyUI, depending on when Comfy updates itstransformersintegration. The same change could also benefit non-training “model surgery” tools such as Heretic, since they may be able to operate on GGUF-backed models without custom conversion or loading paths. -
One practical evaluation use case mentioned is easier swapping between different GGUF quantizations inside the same
transformers-based workflow to compare behavior, such as long-conversation character retention in roleplay chats, without additional loader-specific setup.
-
There is discussion about downstream tooling impact: native GGUF loading in
-
A commenter highlights the main technical implication: because frameworks like Unsloth and Axolotl are built on
Less Technical AI Subreddit Recap
/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo
1. Opus 5.5 Agentic Creative Builds
- Made entirely with Opus 5.5 + $3.21 of OpenRouter API usage (opens in new tab) (Activity: 2308): OP reports a true one-shot autonomous Claude Code generation using Opus 5.5 to create a
30s–60spure-JavaScript whimsical hand-drawn collage animation on “what is the purpose of life?”, including script, assets, animation, concept, and TTS. The run took ~1h20m, cost about$20of Opus usage or ~10%of a Max 5-hour quota, plus$3.21on OpenRouter across8APIs—mostly NanoBanana 2, TTS, and minor auxiliary calls—under a$10OpenRouter budget; OP compares it to an earlier similar post here (opens in new tab). The hosted video link was not accessible during fetch because Reddit returned 403 Forbidden for v.redd.it/cdejwwaqobrh1 (opens in new tab), requiring login/developer-token access. Comments were light on technical critique: one commenter was impressed by the AI-generated voice and framed the result as evidence that creative workers are increasingly exposed to automation, while another expressed concern that this kind of low-cost generated media could flood YouTube feeds. - Jaw literally dropped. I ran the prompt from the “Made entirely with Opus 5.5” post on my own project. Here’s what Claude Code made on its own for about $4. (opens in new tab) (Activity: 1490): A user replicated a prior “Made entirely with Opus 5.5” workflow by giving Claude Code an OpenRouter (opens in new tab) API key capped at
$10and prompting it to autonomously produce a30–60sexplainer video for Friendr.nl (opens in new tab). In ~1.5–2hand for ~$4, it reportedly generated the script/concept, collage-style assets, TTS voice-over, music/SFX, a pure JavaScript canvas animation rendered to MP4, beat-synced animation to narration, and used another model for self-review; an English version took ~30minmore. A commenter reproduced the pattern for “blueprintr” with a similar prompt targeting a45–60sJS/vellum-style animation, noting only minor manual corrections and sharing a Streamable result (opens in new tab). Commenters characterized the result as near-term disruptive for automated video production—e.g. joking that Pixar could soon prompt “make Toy Story 6” —but the thread contained little substantive technical critique beyond anecdotal confirmation that the workflow also worked on another project.-
A commenter shared the exact autonomous generation prompt used to create a
45–60spure JavaScript animated explainer locally runnable in Firefox, with constraints to generate the script, assets, animation, concept, and audio end-to-end. The workflow explicitly allowed Claude Code to use internet resources and a.envOpenRouter API key for a high-quality TTS model, with a max OpenRouter spend of$10; the commenter said only minor corrections were needed and linked the resulting video: https://streamable.com/tsn19a (opens in new tab)
-
A commenter shared the exact autonomous generation prompt used to create a
- Opus 5.5 is insane at making videos (opens in new tab) (Activity: 1329): The post claims Claude Opus 5.5 generated an SNES-style video-game combat video entirely from code, including character assets, animation/timing, fight sequencing, and music, without user-provided assets. The prompt theme was Sydney—Microsoft’s early GPT-4-powered Bing Chat persona with different RLHF behavior, referenced via the archived NYT Bing/Sydney transcript (opens in new tab) —facing Sam Altman and then Claude itself; the Reddit-hosted video could not be independently inspected because
v.redd.it/ghsiido07erh1returned 403 Forbidden. Top comments were uniformly impressed, specifically highlighting the generated video’s timing and pacing as unexpectedly strong; no substantive technical debate or critique was present.-
Commenters highlighted Opus 5.5 as showing unusually strong video-composition behavior, especially around timing and pacing: one noted its “sense of timing and pacing is actually good”. Another compared it to the launch-day viral
p(doom)video, saying outputs are “packed with quick jokes and small details,” suggesting improved scene-level coherence and comedic beat placement rather than just visual generation quality.
-
Commenters highlighted Opus 5.5 as showing unusually strong video-composition behavior, especially around timing and pacing: one noted its “sense of timing and pacing is actually good”. Another compared it to the launch-day viral
- This interactive island was built in 8 hours with Opus 5.5 (opens in new tab) (Activity: 1125): Dan Greenheck built the browser-based interactive island demo TideWater (opens in new tab) in roughly
8 hoursusing Opus 5.5, reportedly relying on simple iterative prompts like “add X” and “make it better” (tweet (opens in new tab)). The demo includes multiple interactive/simulated elements—birds, crabs, fish/whale behavior, wind effects, night lighting, walking/interaction, and boat sailing—and consumed about$1,874.40in tokens, or59%of a Max20xweekly allowance. Commenters were mostly impressed by the scope of the demo beyond the video preview, with one predicting this style of AI-assisted generation could enable “great GTA offshoots” soon. Other reactions were brief/speculative, including jokes about “Opus 50” and one negative comparison that it “looks like crisis.”-
Commenters noted that the demo’s technical scope is clearer when run interactively rather than viewed as a video: users can walk around, interact with objects, and sail the boat, suggesting the Opus 5.5-generated environment includes basic game-loop mechanics beyond static scene generation.
- Several comparisons framed the output as resembling early Crytek / Far Cry 1-era engine visuals, while another commenter specifically highlighted the water physics as visually competitive with some modern AAA titles, though these observations were qualitative rather than benchmarked.
-
Commenters noted that the demo’s technical scope is clearer when run interactively rather than viewed as a video: users can walk around, interact with objects, and sail the boat, suggesting the Opus 5.5-generated environment includes basic game-loop mechanics beyond static scene generation.
2. Claude-Discovered CRISPR-like Enzyme System
- Claude discovered a novel enzyme system with properties reminiscent of CRISPR (opens in new tab) (Activity: 1100): Anthropic reports (opens in new tab) that Claude-agent genome-mining workflows identified a previously uncharacterized bacteriophage system dubbed array-associated reverse transcriptases (ART): an RT gene plus accessory gene adjacent to a long CRISPR-like tandem repeat array. In the described campaign, ~
950Claude agents used210Mtokens over21hours to collect>200kreverse transcriptases, nominate3,500candidate systems, and prioritize20reports; early BSL-1/2 validation found the ART array is transcribed into distinct short RNAs, but Anthropic explicitly says the system’s biological function and any programmable editing utility remain unknown. Commenters were cautiously optimistic, framing this less as an AlphaFold-scale biology result and more as evidence that LLM agents can contribute to original hypothesis generation: “Claude selected an unusual candidate… and brought it to human researchers for validation.” Others speculated that Anthropic’s bio lab could improve public support if it leads to disease-relevant discoveries, while emphasizing that ART is not yet demonstrated to cut/copy/paste DNA or enable gene editing.-
Several commenters emphasized that the reported ART system is not yet comparable to AlphaFold 2 or CRISPR-level functional discovery: Anthropic reportedly shows that the repeat array is transcribed into distinct short RNAs, but the biological function remains unknown and there is no evidence yet of programmable gene editing or a demonstrated mechanism analogous to CRISPR.
- A technical critique argued the work appears incomplete because identifying repeat arrays and showing they produce short RNAs is a fairly standard genomics workflow, with similar analyses already seen in systems such as VIPR. The commenter noted that repeat arrays are already known to be interesting motifs, so the novelty would need to come from either a new biological function or a substantially novel discovery process, neither of which they felt was clearly established.
- One substantive point was that the most important result may be methodological rather than biological: Claude reportedly selected an unusual candidate, noticed an overlooked pattern, assessed novelty, and escalated it for human experimental validation. Commenters framed this as early evidence of AI acting as a research collaborator, even if the enzyme system’s actual importance remains uncertain.
-
Several commenters emphasized that the reported ART system is not yet comparable to AlphaFold 2 or CRISPR-level functional discovery: Anthropic reportedly shows that the repeat array is transcribed into distinct short RNAs, but the biological function remains unknown and there is no evidence yet of programmable gene editing or a demonstrated mechanism analogous to CRISPR.
- The moment Claude agents discover a new molecular mechanism, talking as if they were human, using interjections and cues (opens in new tab) (Activity: 1056): The image (opens in new tab) appears to show Claude agents reasoning through genomic sequence flanks and identifying repeated DNA motifs, with a highlighted realization that the structure may resemble a CRISPR-like or msDNA/retron-like repeat array. The technical significance is not a validated discovery from the screenshot alone, but rather an example of LLM-style agentic hypothesis generation in molecular biology: comparing tandem repeats, spacer regions, and known mobile genetic element architectures such as CRISPR arrays, diversity-generating retroelements, msDNA, and retrons. Comments mostly frame the screenshot as evidence of rapid AI progress, with one user analogizing it to recent gains in mathematics and asking whether “Biology [will be] solved soon?” Others focus on the model’s human-like enthusiasm rather than the biological claim itself.
- Claude Code with Opus 5.5 was reported to generate a 30–60-second pure-JavaScript explainer end-to-end—including concept, script, assets, animation, and TTS—in a one-shot run of about 1 hour 20 minutes; the creator reported about $20 in Opus usage plus $3.21 across eight OpenRouter API calls. A second user reproduced the workflow for Friendr.nl in about 1.5–2 hours for roughly $4, adding music/SFX, narration-synced animation, MP4 rendering, and another-model review; the run used a $10-capped OpenRouter key and reportedly needed only minor corrections. A separate motion-design prompt template starts by asking for 8–12 UI states for a shape to transform into (for example, a button, loader, or player); its creator said the video was entirely code, with no After Effects.
-
Opus 5.5 was also used to build TideWater, a browser-based interactive island, in about eight hours with iterative prompts such as
add Xandmake it better; the reported token cost was $1,874.40, or 59% of a Max 20x weekly allowance. The demo includes walking around, interacting with objects, and sailing a boat, so testing it interactively reveals capabilities beyond a video preview. - For agent builders, LangChain Managed Deep Agents 0.8 adds user/agent memory with access policies, HTTP channels, sandbox file APIs, proxy-authenticated sandboxes, and Parallel web search; its smithtune CLI turns traces into post-training datasets, while Trajectories handles deferred tool calls and context compaction. Databricks also reports that engineers stopped reaching for closed models once open-source models were routed to its internal coding agents.