We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Put hard spend caps on anything your agents deploy
Simon Willison argues that coding agents, and personal agents (which he calls "coding agents wrapped in a less threatening UI"), make it easy to spin up code that costs money: paid API calls, hosted apps, and storage or compute that bills as it grows . A warning email is not enough. Usage-based services need hard caps that cut the service off and return errors. Removing the cap should be an explicit opt-in . His reasoning: most people would rather see errors than a surprise bill of $10,000 or more .
Options you can use now:
- AWS announced monthly project spend limits on Sept 16. A project that hits its limit is paused for the rest of the month. AWS's docs say the feature is still going out to a limited number of customers .
- Google Cloud launched Spend Caps in July. They set a monthly cap on specific services within a project .
Willison also wants agents to lean toward recommending providers with hard caps, and to warn inexperienced builders before they deploy to uncapped services . You can do this today by adding that rule to your agent instructions.
The harness can matter as much as the model
Two results summarized by AINews point the same way. Hugging Face reports that the same model weights scored 62% in one harness and 33% in another. Its multi-harness RL setup uses a proxy that records token IDs and logprobs without changing the harnesses. It lifted LFM2.5-2.6B from 42% to 54% across four harnesses with 31% fewer tool calls, and the trainer, data and seven models are open . Separately, Meta Superintelligence Labs reports that a dedicated controller raised GPT-5.5 on ProgramBench from 63.7% to 71.5%, with the same workers and budget, against 58.0% for Codex . In practice, a model comparison means little unless you know which harness ran it.
T3 Code: queue prompts until limits reset, see spend across tools
- Queued messages: T3 Code Nightly now lets you queue messages that fire when your usage limits reset .
- Usage view: Theo improved the view so you can see where your "spend" goes and how models behave on your own data. It counts all Claude Code and Codex usage on your machines, not just usage inside T3 Code .
- Model mix: Opus 5.5 is the first model to pass 50% of T3 Code traffic; "literally half of all prompts go to Opus" .
- Orchestrator V2 is in the latest Nightly, and Theo is asking for reports on what broke or confused people . The mobile app only works through the TestFlight/beta build (install docs) .
- Users: T3 Code passed 400K users and gained about 20K more within a day .
Claude Code mods in use
Anthropic added a built-in plugin, "You should Know." It scans Claude's output for important information you might miss. Enable it with /plugin enable cc-plugin-you-should-know@builtin. AINews describes it as spinning off a side agent, and describes mods as plugins with middleware-like hooks into Claude Code . A community example: @shawnbuilds built a custom Jev memory harness as a mod and posted the full prompt to rebuild it .
On a related habit, swyx says: "always run some kind of skill review/cutter after every model-created skill" .
Codex: DevDay features and a product refocus
Riley Brown's DevDay walkthrough covers the features most relevant to Codex users:
- Codex CLI: easier voice input, a new agents view and better prompt editing .
- Codex Cloud: you set up an environment once (signed into Clerk, Convex, Vercel, GitHub and so on), then reuse it for tasks run from your phone, with no local machine left on .
- Code review: PRs can be reviewed from a side panel. You may need to pin "code review" there first .
- GPT-6.1 Sol: OpenAI pitches it as a lower-cost model for coding and computer use. Brown quotes $2 in / $10 out per million tokens, against $10 / $50 for GPT-6 Astra .
- Ultra-fast mode: about 8x faster output at 6x the usage, and only on the $500/month plan .
OpenAI's Tibo says the Codex team is "locking in": the only work now underway is simplification, efficiency to allow more usage, groundbreaking features and new models, because feedback says users want things simpler .
Smaller items
- A dot clearing an inbox: Tibo reached inbox zero by giving his dot the goal. It deleted unneeded categories in validated batches, labeled email by type of work, and walked him through replies while it looked up context in the background . Simon Willison, by contrast, isn't sure when to use a dot over ChatGPT, because his dot seems built around a single conversation while he prefers managing context across threads .
- Cheap computer-use targeting: ThePrimeagen uses Cloudflare's decision API for computer use. He raises its probabilities to chosen powers, then applies "probabilistic centering" so a command like "Click the monitor icon in the menu bar" lands on the right icon .
- Pi 1.0 now includes native MCP support in Codemode by default, plus deferred tool loading, Anthropic cache warming and mid-conversation system messages .
- Apple: Apple plans new Mac privacy controls that warn about granting broad data access to third-party software, including AI agents . DHH expects this to make macOS harder to use productively with agents .
- Pi 1.0 added native MCP support to Codemode, deferred tool loading, Anthropic cache warming, mid-conversation system messages, and support for non-LLM/image models and virtual-model extensions. Claude Code mods expose middleware-like hooks; its “You should know” plugin runs a side agent that flags important output the user might miss.
-
T3 Code’s rewrite added cross-provider
delegate_task, thread forking, mid-thread model switching, subagent lineage views, scheduled tasks, Pi support, and an ACP registry. Cursor Rollouts can identify the PR behind a regression, open an issue, and offer a one-click cloud-agent fix. - Agent performance can vary substantially with the harness: identical weights scored 62% in one harness and 33% in another. Multi-harness training raised LFM2.5-2.6B from 42% to 54% across four harnesses and cut tool calls by 31%; the trainer, data, and seven trained models were released openly. Meta Superintelligence Labs also reported that a dedicated controller raised GPT-5.5’s ProgramBench score from 63.7% to 71.5% with the same workers and budget, versus 58.0% for Codex.
-
A user tested Qwen3.8-27B IQ4_XS via
llama.cppon one 24 GB RTX 4090 at a 196,608-token context, reporting 12/12 on a code-review task and 40/43 hidden tests on one DeepSWE task—but the binary pass score was 0, and the author cautioned this did not establish broad parity across the 113-task benchmark.
- Codex Cloud provides a way to run coding tasks without leaving a local computer on: configure a cloud environment with services such as Clerk, Convex, Vercel, and GitHub, reuse it across tasks, and work from a phone.
- Codex CLI gained voice input, an agents view, and improved prompt editing. Codex Code Review can review pull requests from its side panel; the presenter was unsure whether the announced security dashboard was publicly available.
- OpenAI’s Agents API lets developers run agents in their own products; the video says setup requires a platform project and API key, and agents can use an OpenAI-hosted browser for computer use.
- OpenAI’s newly announced 6.1 model is positioned as a lower-cost model for coding and computer use, which the presenter described as a practical workhorse based on initial use. Ultra-fast mode was described as roughly eight times faster while using six times the regular usage, and limited to the $500/month plan.
Treat hard spend caps as a deployment guardrail for coding-agent-built services: agents lower the friction of creating software that can incur paid API, hosting, storage, and compute charges, so the author argues usage services should stop at a hard limit rather than only send warnings, with removing the cap requiring explicit opt-in. AWS announced monthly project spend limits on September 16, 2026, which pause a project when its limit is reached, although the feature was still rolling out to a limited number of customers; Google Cloud had launched monthly Spend Caps for specific services within a project in July. The author suggests agents could recommend providers with hard caps and warn against deploying to uncapped services.
Riley endorsed a principle for AI-assisted software creation: bring your own point of view and decide what deserves to exist; asking AI to choose without your intention risks plausible, easy-to-approve output, while greater speed makes judgment more important, not less.
Claude Code’s “You should Know” plugin scans Claude’s output for important information you might miss, helping keep you informed; enable it with /plugin enable cc-plugin-you-should-know@builtin. Addy Osmani calls it useful .
ThePrimeagen says he uses Cloudflare’s Decision API for computer use, applying what he calls a “noul” approach: raising probabilities to selected powers and then probabilistically centering them to click the correct icon. His example instruction is “Click the monitor icon in the menu bar.”
Ben Tossell says personal agents do not use Opus 5.5 yet . He calls Opus “the goat” with early OpenClaw .
@_justelias says Kody can turn needs that would otherwise require a small service, script, credential, or webhook into capabilities usable by any agent, reducing glue work and avoiding rebuilding your setup for each agent . @kentcdodds endorses its broad applicability .
Geoffrey Huntley argues that the compile/restart/deploy loop could be collapsed, and predicts a shift toward adapting the older “actors/repair the image (don’t rebuild it)” idea for systems designed for LLMs as programmers rather than humans—a conceptual direction for coding-agent system design, not a concrete workflow.
- Codex CLI adds voice input, an agents view, and improved prompt editing. Codex Cloud runs coding tasks without keeping a local computer on; configure an environment with services such as Clerk, Convex, Vercel, and GitHub once, then reuse it for cloud tasks from the phone.
- Codex can review pull requests from its code-review panel. A security dashboard for checking app vulnerabilities was also announced, though the creator was unsure whether it was publicly available yet.
- OpenAI’s Agents API is presented as a way to run Codex-powered agents inside developer-built products, using a platform project and API key, with computer use in an OpenAI-hosted browser. The announced Decisions API instead targets fast, constrained choices such as routing incoming messages to sales, billing, or support; the creator said it was not released yet.
- The video says Ultra-fast mode generates output about eight times faster but uses six times the usage of regular mode and is limited to the $500/month Pro plan.
Run a skill review/cutter after every skill created by a model; the post strongly recommends making this a consistent step in coding-agent workflows.
Jason Zhou recommends a Claude Code Mods tutorial and highlights the /jev-memory mod as “pretty cool.” The linked post’s creator says they built a custom Jev memory harness using @typesafeai via @treg_ai and points readers to a full rebuild prompt.
Peter Steinberger said OpenClaw’s Android app had been in review limbo for over a week and asked whether anyone at Google could help; the post does not specify the review process or the app’s availability.
Riley proposed a game-hosting platform designed for existing coding agents: a user would paste its link into Codex, Claude, or Grok and ask the agent to make a game multiplayer; the agent would use the platform’s docs to host and list it. He emphasized getting real-time database support, latency, and low-friction access right, and making the platform free initially.
Riley Brown says he plugged in a ModRetro and asked Codex whether it was compatible with Muse Gadgets, presenting this as a way to make the device smart.
Theo said contributors to T3 Code had collectively burned “$10m+ in tokens.” He also reported using up four Claude accounts without hitting any of his Fable limits.
T3 Code had surpassed 400,000 users, and Theo later reported another 20,000 users gained since that update.
Riley Brown says coding agents have moved beyond routine SaaS apps into multiplayer game mods, adding agents to older hardware, and designing objects for 3D printing, citing Codex and Claude Code as tools used for these projects.
Tibo says current work is limited to simplifications, greater efficiency as usage grows, breakthrough features, or new models; he says user feedback clearly favors making things simpler.
Theo asked users to try the latest T3 Code nightly and Orchestrator V2, then report what went wrong, what was confusing, what broke, and what to prioritize next. T3 Code also has a mobile app, but users must install its TestFlight/beta version; Theo linked installation instructions.
[AINews] not much happened today
If you’re even seeing this, you should probably just go enjoy your weekend.
AI News for 10/1/2026-10/2/2026. We checked 12 subreddits, 544 Twitters (opens in new tab) and no further Discords. AINews’ website (opens in new tab) lets you search all past issues. As a reminder, AINews is now a section of Latent Space (opens in new tab). You can opt in/out (opens in new tab) of email frequencies!
AI Twitter Recap
GPT-6.1 Sol and Sonnet 5.5 Reshape the Cost–Performance Frontier
- GPT-6.1 Sol launch: OpenAI priced Sol at $2/$10 per million input/output tokens, compared with $10/$50 for Astra (pricing summary (opens in new tab)).
- Claimed results: It reportedly beats GPT-6 Sol by 6.4 points on DeepSWE v1.1 and Opus 5.5 by 2.2 points on AutomationBench (summary (opens in new tab)).
- Positioning: OpenAI staff describe it as “good, cheap AND fast” (@reach_vb (opens in new tab)).
- Codex usage: A global Codex usage reset was set for Oct 2 at 10AM PT (@reach_vb (opens in new tab)).
- Tool use: Sol reportedly “REALLY loves codemode,” consistent with GPT models being trained on it (@badlogicgames (opens in new tab), codemode note (opens in new tab)).
- Claimed results: It reportedly beats GPT-6 Sol by 6.4 points on DeepSWE v1.1 and Opus 5.5 by 2.2 points on AutomationBench (summary (opens in new tab)).
- Agent Arena placements: Sol [Max] entered at #5 (+11.23%) with a $0.56 median cost per task (@arena (opens in new tab)).
- Sol cost comparison: That is 39% cheaper than GPT-6 Sol while scoring 1.52 points higher. It is 81% cheaper than Astra while landing within 1.04 points.
- Sonnet 5.5: Sonnet 5.5 [Max] debuted at #3 (+12.5%) and ranked #1 in the Chat category. It costs $2.74 per task, versus $1.58 for #2 Opus 5.5, which keeps it off the Pareto frontier (debut (opens in new tab), frontier (opens in new tab)).
- Anthropic’s position: Anthropic models now hold the top three Agent Arena spots.
- Sol cost comparison: That is 39% cheaper than GPT-6 Sol while scoring 1.52 points higher. It is 81% cheaper than Astra while landing within 1.04 points.
- Code and Text Arena: Sol briefly entered WebDev at #3 before Sonnet 5.5 pushed it to #4 (weekly recap (opens in new tab)).
- Sonnet on WebDev: Sonnet now sits 2 points behind GPT-6 Astra [Max] at 80% lower cost.
- Gemini 4 Argon: Argon [High] took #1 in Text Arena.
- Open models: MiMo-V2.6-Pro and Flash entered Agent Arena at #5 and #9 among open models.
- Sonnet on WebDev: Sonnet now sits 2 points behind GPT-6 Astra [Max] at 80% lower cost.
- Other independent evals: WeirdML v3 finds Sol very token-efficient, close to Astra but with a lower peak. On the same benchmark, Sonnet 5.5 beats Opus 5 and Grok 4.7 beats Kimi-K3; these results are incomplete (@htihle (opens in new tab)).
- Reasoning style: Design Arena read 324 thinking summaries. It found that Astra hedges about 20× as often as Opus 5.5, while Opus commits early in about 4 of 5 summaries (@DesignArena (opens in new tab)).
- Step 5 Preview: StepFun’s model ranks #7 among open-weight models on Vals at $2.54 per task. It averages nearly two hours per task and has a 1M-token context window (Vals (opens in new tab), details (opens in new tab)).
- Reasoning style: Design Arena read 324 thinking summaries. It found that Astra hedges about 20× as often as Opus 5.5, while Opus commits early in about 4 of 5 summaries (@DesignArena (opens in new tab)).
- Rumors (unconfirmed):
- Fable 5.5: Claude Fable 5.5 is rumored for next week and said to outperform an “Astra 6.1” that was reportedly delayed over security concerns. The poster says he cannot verify either claim (@kimmonismus (opens in new tab), follow-up (opens in new tab)).
- GPT-6 Astra Lite: A “GPT-6 Astra Lite” listing has been spotted, which @scaling01 speculates is the same model as Sol (@scaling01 (opens in new tab)).
- Fable 5.5: Claude Fable 5.5 is rumored for next week and said to outperform an “Astra 6.1” that was reportedly delayed over security concerns. The poster says he cannot verify either claim (@kimmonismus (opens in new tab), follow-up (opens in new tab)).
- Decision models and open weights: llama.cpp added a
/v1/systemoneendpoint for local “Jev-style” decision-model inference (@ggerganov (opens in new tab)).- Running locally: Models are launched with
llama serve -hf ggml-org/Kev-4B-GGUF(@ClementDelangue (opens in new tab)). Jared Palmer published a post on how Kev 1.0 works (post (opens in new tab)).- Ecosystem: Perplexity claims pplx-decider-v1-27b averages 85.7% across 11 benchmarks, ahead of Jev (@AravSrinivas (opens in new tab)). Clef decision models are now on Ollama (@lucataco (opens in new tab)).
- Skeptical view: @mervenoyann calls decision models a rebrand of zero-shot classifiers (tweet (opens in new tab)).
- Calibration analysis: A blog post links Jev-style calibration to value and Q-function prediction (@SOURADIPCHAKR18 (opens in new tab)).
- webAI TwIL-LM3-Pro: This 3.66B model is post-trained from Granite 4.2. In webAI’s tests it roughly matches Qwen3-8B on formal logic. The Q4 GGUF is 2.09 GiB and the license is non-commercial (@kimmonismus (opens in new tab)).
- Reka RIDM: Reka released an inverse dynamics model under Apache 2.0. It is trained on games, generalizes to real video and extracts motor and camera actions (@RekaAILabs (opens in new tab)).
- Running locally: Models are launched with
Agent Harnesses, Assistants and Developer Tooling
- OpenAI dots: Sam Altman calls dot his favorite OpenAI product, saying it improves daily as it learns his workflow (@sama (opens in new tab)).
- Capabilities: Dot keeps context across apps, coordinates Codex tasks and flags items that need attention (@OpenAIDevs (opens in new tab)).
- Comparisons: One user prefers Grokbot’s multi-agent “chief of staff” setup (@kimmonismus (opens in new tab)). A DIY clone uses Pi, a Telegram gateway and any model (@_alejandroao (opens in new tab)).
- Capabilities: Dot keeps context across apps, coordinates Codex tasks and flags items that need attention (@OpenAIDevs (opens in new tab)).
- Muse Gadgets: Meta open-sourced ESP32 firmware and a Linux SDK for building hardware that works with Muse (@natfriedman (opens in new tab)).
- Muse Home Link: Meta made 5,000 units of its own smart-home bridge, free for subscribers while supplies last (@alexandr_wang (opens in new tab), shipping (opens in new tab)).
- Extensible harnesses: DeepSeek Harness shipped desktop builds for macOS and Windows; Linux users install
@deepseek-ai/dshfrom npm (@deepseek_ai (opens in new tab)).- Claude Code mods: Mods are plugins with middleware-like hooks into Claude Code (@lydiahallie (opens in new tab)). The new “You should know” plugin spins off a side agent that flags important output the user might miss (@ClaudeDevs (opens in new tab)).
- Pi Durable: Pi now runs on Cloudflare Durable Objects via agents SDK v0.26.0, alongside Pi’s v1.0 release (@mattzcarey (opens in new tab), @badlogicgames (opens in new tab)).
- Context: @omarsar0 frames these releases as a shift toward malleable harnesses (thread (opens in new tab)).
- Claude Code mods: Mods are plugins with middleware-like hooks into Claude Code (@lydiahallie (opens in new tab)). The new “You should know” plugin spins off a side agent that flags important output the user might miss (@ClaudeDevs (opens in new tab)).
- T3 Code orchestrator rewrite: The project passed 400K users (@theo (opens in new tab)). Its 4-month PR, with 823 commits across 1,912 files, has now merged (@maria_rcks (opens in new tab)).
- New features: The rewrite adds Pi support, cross-provider
delegate_task, an ACP registry, thread forking, mid-thread model switching, subagent lineage views and scheduled tasks (feature list (opens in new tab)).
- New features: The rewrite adds Pi support, cross-provider
- Platform updates: OpenAI’s Agents API added one-call browser computer use, Bedrock Managed Agents and portable environments. It also claims 99.97% turn reliability and 20% faster tool calls (@stevendcoffey (opens in new tab)).
- Cursor Rollouts: When Rollouts catches a regression, it finds the offending PR, opens an issue and offers a one-click cloud agent fix (@cursor_ai (opens in new tab)).
- Cloudflare: Sandbox SDK 1.0 gives Durable Objects direct control over sandbox containers (@CFchangelog (opens in new tab)). Cloudflare also launched request Traces (@WalshyDev (opens in new tab)).
- Cursor Rollouts: When Rollouts catches a regression, it finds the offending PR, opens an issue and offers a one-click cloud agent fix (@cursor_ai (opens in new tab)).
Research: Agent Training, Long-Horizon Control and AI for Math
- Multi-harness RL (Hugging Face): The same model weights score 62% in one harness and 33% in another (@huggingface (opens in new tab)).
- Method: A proxy speaks the OpenAI, Anthropic and Gemini API formats and records sampled token IDs and logprobs for training, with no changes to the harnesses themselves.
- Results: LFM2.5-2.6B improved from 42% to 54% across four harnesses and made 31% fewer tool calls. SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.
- Release: The trainer, data and all seven trained models are open.
- Method: A proxy speaks the OpenAI, Anthropic and Gemini API formats and records sampled token IDs and logprobs for training, with no changes to the harnesses themselves.
- Credit assignment and RL efficiency: ProVer has a judge locate the decisive trajectory segment, then uses rollouts on either side to set that segment’s advantage. It reports +9.91% (Qwen3.5-2B) and +7.12% (Qwen3.5-4B) relative gains over GRPO (@omarsar0 (opens in new tab)).
- Partial rollouts: AC2 uses a learned critic to score token chunks, so training needs only partial rollouts (@wen_kaiyue (opens in new tab)).
- Frontier Learning: The method targets problems at the edge of capability, since problems a model always or never solves give zero GRPO gradient (@robinfaro13 (opens in new tab)).
- Sharpening Tax: The paper quantifies the loss of pass@K scalability after post-training and proposes PTGS, a per-prompt temperature sampler (@iScienceLuvr (opens in new tab)).
- SFT vs RL: Another paper finds SFT generalizes worse because its data is off-policy, not because of the objective. Rewriting expert trajectories in the base model’s style closes the gap (@maximelabonne (opens in new tab)).
- Partial rollouts: AC2 uses a learned critic to score token chunks, so training needs only partial rollouts (@wen_kaiyue (opens in new tab)).
- Long-horizon control and context: Meta Superintelligence Labs reports that a dedicated controller lifts GPT-5.5 on ProgramBench from 63.7% to 71.5%, using the same workers and budget, versus 58.0% for Codex (@dair_ai (opens in new tab)).
- Context compression: Microsoft’s training-free FOCUS cuts peak context by up to 48% and raises task success by up to 8.9 points (@dair_ai (opens in new tab)).
- Long-context degradation: NVIDIA’s Long-Transduction study measures a 62.8% accuracy drop from 4K to 128K context across seven open models (@dair_ai (opens in new tab)).
- Multi-agent coordination: In AgentWorld, fewer than a third of multi-agent actions help complete the task, and coordination tasks reach only 12% success (@omarsar0 (opens in new tab)).
- Apple LoopCD: The method halves recurrent loops while raising AIME 2024 pass@1 from 61.88% to 73.33% (@arankomatsuzaki (opens in new tab)).
- Context compression: Microsoft’s training-free FOCUS cuts peak context by up to 48% and raises task success by up to 8.9 points (@dair_ai (opens in new tab)).
- AI on open math problems: Meta released six papers on open problems produced with Muse Spark 1.1 and 1.2 through plain meta.ai chat, with no custom scaffold (@AIatMeta (opens in new tab), list (opens in new tab)).
- Process: Each paper labels which passages were drafted primarily by humans or by AI, and a second group of mathematicians reviewed the work.
- Google Cogentic: This Gemini multi-agent system produced new results on five open theory problems (@omarsar0 (opens in new tab)).
- Cogentic design: Each draft must pass two adversarial verifiers, and agents share a ledger of verified lemmas. Most problems took about 100 calls; the hardest took about 1,000.
- Process: Each paper labels which passages were drafted primarily by humans or by AI, and a second group of mathematicians reviewed the work.
- Image post-training: Arena combined a Bradley-Terry reward model with faithfulness, constraint and anti-reward-hacking rewards (@arena (opens in new tab)).
- Results: FLUX.2-dev gained 69 Elo to 1202, and Ideogram 4 gained 20 Elo to 1224.
Benchmarks, Eval Integrity and Safety
- Research-taste benchmarks: ScholarCatalyst asks agents to find the “catalyst papers” behind research projects. It is labeled by 184 lead authors on 207 of their own projects and is described as far from saturated (@yoonholeee (opens in new tab)).
- EurekaBench: This benchmark tests whether agents can discover genuinely new insights across six science domains (@JiayiiGeng (opens in new tab)).
- Vals Web Search Index: The index holds model and harness constant, swaps only the search tool, and scores final answers on finance and legal tasks (@ValsAI (opens in new tab)).
- Validation: Agents score 2.9% (legal) and 7.4% (finance) without search, versus 30–50% with it. Vals also cites a study in which a model answered 44.5% of BrowseComp without search (details (opens in new tab)).
- SWE bug-finding bench: In this new benchmark, agents start from an older commit and are scored against real bugs fixed in later commits.
- Critique: Lucas Beyer argues it mainly tests recall and that the construction is easy to train toward (@giffmana (opens in new tab)).
- Authors’ response: The authors say training for bug-finding is fine as long as the test set is excluded (@OfirPress (opens in new tab)).
- Critique: Lucas Beyer argues it mainly tests recall and that the construction is easy to train toward (@giffmana (opens in new tab)).
- Eval integrity question: David Rein asks whether Harbor, the framework behind Terminal Bench, lets agents modify their trajectories before evaluation. He notes he may be misreading the code (@idavidrein (opens in new tab)).
- Offensive capability of open models: The Batch reports GLM-5.3 nearly matched Claude Mythos at exploiting vulnerabilities, 12% vs 14% (@DeepLearningAI (opens in new tab)).
- Disputed claim: One commentator says GLM-5.3 Flash exceeds Mythos Preview on ExploitBench (@teortaxesTex (opens in new tab)).
- Uncensored variant: An uncensored GLM-5.3 is circulating on Hugging Face (@kimmonismus (opens in new tab)).
- Disputed claim: One commentator says GLM-5.3 Flash exceeds Mythos Preview on ExploitBench (@teortaxesTex (opens in new tab)).
- Safety research and safeguards: A new paper proposes using internal signals during training to improve alignment without degrading white-box monitoring (@lenalibon (opens in new tab)).
- NeurIPS acceptance: “Models That Know How Evaluations Are Designed Score Safer” was accepted at NeurIPS 2026 (@HaritzPuerto (opens in new tab)).
- False positives: Opus 5.5 frequently triggers “reasoning extraction” safeguards during spectrogram syllable labeling (@ChaseBrowe32432 (opens in new tab)).
- NeurIPS acceptance: “Models That Know How Evaluations Are Designed Score Safer” was accepted at NeurIPS 2026 (@HaritzPuerto (opens in new tab)).
- Emergent world knowledge: Asking a model “land or water?” for 16,200 lat/long coordinates and plotting the answers yields a recognizable world map (@karpathy (opens in new tab)).
Inference, Hardware and Systems
- Ascend 950 via DeepSeek kernels: An analysis of DeepSeek’s open-sourced DeepGEMM, FlashMLA, TileKernels and DeepEP infers the chip’s layout (@ZhihuFrontier (opens in new tab)).
- Estimated specs: The chip has 32 AI cores, each pairing one Cube core with two Vector cores. Estimated peaks are about 432/865/1,730 TFLOPS in BF16/FP8/FP4.
- Capacity: Supply may be limited, despite claims that 950s went on sale in August (@teortaxesTex (opens in new tab)).
- Estimated specs: The chip has 32 AI cores, each pairing one Cube core with two Vector cores. Estimated peaks are about 432/865/1,730 TFLOPS in BF16/FP8/FP4.
- Prime Inference: Prime Intellect stores the MLA latent in NVFP4, shrinking rows from 576 to 352 bytes and fitting about 50% more cached tokens than FP8 (@PrimeIntellect (opens in new tab)).
- Stack: It serves GLM-5.3 on vLLM and Dynamo, and the sparse-MLA kernel is going to FlashInfer (@vllm_project (opens in new tab)).
- Low-precision benchmarking: Stas Bekman measured NVFP4 about 9% more efficient than MXFP4 on B200, with higher accuracy (@StasBekman (opens in new tab)).
- mamf-finder: The tool now benchmarks FP8, MXFP8, MXFP4 and NVFP4 (update (opens in new tab)).
- Memory and speed: NVHBM moves the memory controller into a custom base die, claiming up to 30% more bandwidth and 15% lower power than HBM4E (@vikramskr (opens in new tab)).
- Disputed economics: Micron says NVHBM will improve its margins; @vikramskr disputes this (counterpoint (opens in new tab)).
- Volantis: The startup is targeting up to 10K tokens/s per user on models over 10T parameters using optics (@omarsar0 (opens in new tab)).
- Cerebras: Altman called Cerebras a close partner on speed (@sama (opens in new tab)).
- Disputed economics: Micron says NVHBM will improve its margins; @vikramskr disputes this (counterpoint (opens in new tab)).
- Capacity economics (Epoch): Epoch estimates AI infrastructure could soon support hundreds of millions to billions of agents (@EpochAIResearch (opens in new tab)).
- Demand gap: Just 20% utilization implies $2.6–5.3T in annual spending, against roughly $1T in lab revenue by the end of 2027 (details (opens in new tab)).
- Platforms: SemiAnalysis rates Google’s GPU clusters Gold tier and notes the ConnectX NCCL plugin now auto-activates (@SemiAnalysis_ (opens in new tab)).
- Federated learning: Google Research launched TEE-backed federated learning with verifiable differential privacy (@GoogleResearch (opens in new tab)).
Industry and Policy
- Anthropic and the Vatican: The NYT reports that Chris Olah raised pulling out of the Pope’s AI encyclical launch, whose text rejects machine consciousness (@ChristopherHale (opens in new tab)).
- Lobbying: Olah’s team reportedly lobbied the Pope’s advisers to take model consciousness seriously. He ultimately attended (@kimmonismus (opens in new tab)).
- Context: The article opens with Olah saying “we don’t know if A.I. models are conscious” (@buccocapital (opens in new tab)).
- Criticism: Aidan Gomez criticized the campaign as moral arrogance (@aidangomez (opens in new tab)). Lucas Beyer noted a transcript wording change from “create” to “train” (@giffmana (opens in new tab)).
- Lobbying: Olah’s team reportedly lobbied the Pope’s advisers to take model consciousness seriously. He ultimately attended (@kimmonismus (opens in new tab)).
- Anti-safety influence campaign: A report describes a group planning to spend at least $100M, run by a former White House deputy chief of staff, that frames AI warnings as a coordinated campaign (@NeelNanda5 (opens in new tab)).
- New organizations: Nathan Lambert and Tom Zick launched Trillium Labs, a non-profit for open post-training recipes and infrastructure (@natolambert (opens in new tab)).
- Funding: Initial support comes from Halcyon Futures and Schmidt Sciences.
- Underdog: The private on-device AI startup announced backing from a16z, Khosla and others (@0xSigil (opens in new tab)).
- Funding: Initial support comes from Halcyon Futures and Schmidt Sciences.
- Governance and markets: Yoshua Bengio joined Canada’s new National Council on AI (@Yoshua_Bengio (opens in new tab)).
- Meta: Meta has parted ways with Virtue AI (@AndrewCurran_ (opens in new tab)).
- Nvidia: Bloomberg reports a record high near $5.7T market value after a $150B buyback increase (@kimmonismus (opens in new tab)).
- Meta: Meta has parted ways with Virtue AI (@AndrewCurran_ (opens in new tab)).
Top tweets (by engagement)
- NYT report on Olah and the Pope’s encyclical (opens in new tab) — 28.7K
- Karpathy’s “land or water” eval (opens in new tab) — 16.4K
- Claude Code “You should know” plugin (opens in new tab) — 7.3K
- Altman on dots (opens in new tab) — 6.8K
- Muse Gadgets announcement (opens in new tab) — 4.9K
- DeepSeek Harness desktop builds (opens in new tab) — 4.3K
- Altman on the Cerebras partnership (opens in new tab) — 4.1K
- Trillium Labs launch (opens in new tab) — 2.6K
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. Qwen Local Inference: 27B Benchmarks, Fine-Tunes, and MTP
- I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. (opens in new tab) (Activity: 1192): OP built backburner, a
llama.cppfork / distributed inference setup (opens in new tab) that offloads part of Qwen 3.8 27B IQ4_XS from a24 GBM4 Pro MacBook to an iPhone 17 Pro Max over10 Gb/sUSB-C: the Mac runs layers1–40, streams activations, and the phone runs layers41–64using Metal 4 tensor ops. Reported end-to-end prefill gains vs Mac-only were+35%at8k,+44%at16k,+29%at32k, and+30%at48k; a cold27ksession improved from245 sstockllama.cpp/228 sfork Mac-only to168 swith the phone. Above64kcontext, the phone instead hosts old KV pages—up to roughly5.7 GB, enabling196k–229k8-bit context allocation—and computes old-key attention, with a140kcontext test improving generation latency from279 ms/tokento176 ms/tokenwhen adding Neural Engine-compiled16kkey pages.- A technically relevant follow-up asked whether the same iPhone-as-secondary-GPU approach could extend to iPads, especially higher-end iPad Pro configurations with more capable Apple Silicon and potentially more RAM. The implication is that iPads might provide better offload performance or hold a larger portion of the context window than an iPhone, making them a stronger companion device for local LLM inference.
- Qwen3.8-27B-Humanlike-Chat 2.0: texts like a human, now with tool calls and better instruction following (opens in new tab) (Activity: 805): LessThanThreeAI released Qwen3.8-27B-Humanlike-Chat 2.0, a merged LoRA over huihui-ai’s abliterated Qwen3.8-27B, available as GGUF/BF16/LoRA on Hugging Face (opens in new tab) with a demo Space (opens in new tab). v2 replaces plain SFT with on-policy distillation: the student generates replies while two teachers score tokens—v1 + hidden “text like a person” instruction for chat/character behavior, and the base model for instruction-following, tools, and code—improving tool-use and controllability while preserving informal texting style. Reported evals vs the abliterated base: IFBench
37.3 → 43.7, When2Call48 → 58, BFCL irrelevance60 → 78, ties/slight gains on IFEval/GSM8K/BFCL simple (83.5 / 89.1 / 98), but regressions on MMLU-Pro (78.5 → 72.5) and LiveCodeBench (56 → 51); a custom “ishuman” judge benchmark rated it as human-written23.5%vs0.3%for the abliterated base and15.1%for official Qwen3.8-27B. Technical discussion in the top comments was sparse; the only relevant critique was that the model’s “humanlike” register may read more like teenage texting than broadly human conversation.- A commenter raised a model-transfer question: whether the same humanlike chat fine-tuning/alignment method used for Qwen3.8-27B-Humanlike-Chat 2.0 would produce similar results on Gemma 4 31B. This is the only technically substantive thread, touching on cross-architecture generalization of the training recipe and whether behavior-style tuning would carry over to a larger Gemma-family model.
- The gap is smaller than they told you: local 27B nearly matches frontier on real code tests (opens in new tab) (Activity: 730): OP reports a single-task DeepSWE/local-code benchmark run using Qwen3.8-27B GGUF via llama.cpp b11115 + llama-swap v257 on 1× RTX 4090 24GB, specifically
Qwen3.8-27B-UD-IQ4_XS.gguf(14.25GB) from unsloth/Qwen3.8-27B-GGUF (opens in new tab), atctx-size 196608,IQ4_XS,q8_0K/V cache, speculative MTP draft, and DeepSeek-style reasoning budget4096. Measured results:115 tok/sdecode,22,934 MiBpeak VRAM,12/12on a code-review task, and on one DeepSWE task40/43hidden tests plus109/109existing tests, i.e. partial0.980but binary pass0; OP later corrected the comparison: the cited96.6%was mean partial across all published trials, while the frontier subset for that task was99.8%partial and85.3%pass, from DeepSWE v1.1 raw data (opens in new tab). The linked writeups cover the 24GB fit/context setup (context ceiling (opens in new tab)) and the task-level DeepSWE result (local confidence (opens in new tab)); OP emphasizes this is task-specific, not a claim that a 27B local model matches frontier models broadly across the113-task benchmark. Commenters were skeptical of the broader framing: one user with both “Flash and 27B” said “the gap is real,” and another argued the conclusion is wrong because even frontier coding models are uneven and~24–72Bmodels may handle discrete subtasks but often lose value once humans must decompose larger engineering work into model-sized tasks.-
Several commenters argued the claimed near-parity is likely an artifact of a saturated benchmark: a local
27Bmodel, even atQ8, can perform well on small/discrete coding tasks but still fails on harder real-world tasks requiring frontier models such as Claude Opus.-
A recurring technical objection was that coding evaluations often underweight project-level decomposition:
24B–72Blocal models may solve isolated tickets, but for larger work items the human effort needed to break problems into model-sized subtasks can exceed the productivity gains. -
Users with hands-on experience running both Gemini Flash and local
27Bmodels reported that the performance gap remains substantial, especially for nontrivial coding workloads where frontier models provide better reliability and task completion.
-
A recurring technical objection was that coding evaluations often underweight project-level decomposition:
-
Several commenters argued the claimed near-parity is likely an artifact of a saturated benchmark: a local
- Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp (opens in new tab) (Activity: 405):
llama.cppPR #29761 (opens in new tab) adds MTP speculative decoding support for Qwen3.8-Flash Next via--spec-type draft-mtp, merged into theaman/qwen4-optbranch after ~17hof development. Reported DGX Spark benchmarks for Qwen3.8-Flash-Nextiq4_xswith-np 1 -lzm on --spec-draft-n-max 3show decode throughput improving from28.36to43.88 tok/s(1.55×), latency speedup of 1.54×, and mean speculative acceptance of0.640across24tasks; GGUF quants are available on Hugging Face (opens in new tab). Commenters noted the model is still impractically large for many local setups: theIQ4_NLGGUF is split into a tiny10.9 MBshard plus a102 GBshard, undercutting the idea of casually switching from Qwen 3.8 27B. One commenter also noted that Gufo supports MTP.-
One commenter reported that enabling MTP made inference slower in their testing, arguing it may be more useful for dense models than very large overall architectures where the MTP head has a low acceptance/hit rate. Their hypothesis is that the MTP head cannot effectively predict/compress enough of the larger model’s behavior, reducing speculative decoding benefit.
- A user noted that Gufo already supports MTP, implying llama.cpp is catching up with existing MTP-capable tooling/backends for Qwen-style experimental models.
- Another commenter said they had been using an EXL3 version through tabbyapi because llama.cpp GGUF inference was “way, way slower” for their workload. They planned to retest after this PR, but their prior experience suggests EXL3/tabbyapi may still be a performance baseline to compare against for Qwen4Exp/MTP support.
-
One commenter reported that enabling MTP made inference slower in their testing, arguing it may be more useful for dense models than very large overall architectures where the MTP head has a low acceptance/hit rate. Their hypothesis is that the MTP head cannot effectively predict/compress enough of the larger model’s behavior, reducing speculative decoding benefit.
2. Local Agent Tooling: Decision Models and MCP
- Pi 1.0 released - MCP support now included by default (opens in new tab) (Activity: 679): Earendil released
Pi 1.0, a stable version of its minimal agent harness, with Codemode now including native MCP support by default plus non-LLM/image model support, virtual-model extensions, deferred tool loading, Anthropic cache warming, mid-conversation system messages, and TUI updates. The release also introduces experimental MIT-licensed Pi Durable for longer-running agentic applications beyond terminal/coding-agent workflows, while retaining Pi’s minimal/extensible architecture. Top comments focused on naming ambiguity— “pi” collides with many AI/dev tools—and requested clarification of what Codemode is. One commenter linked Earendil’s rationale for MCP support: “You said no MCP” (opens in new tab).- A commenter linked the maintainer’s rationale for reversing course on MCP support in Pi 1.0, pointing to the post “You said no MCP” (opens in new tab). The thread notes that MCP is now included natively/by default, after earlier resistance from the creator based on project ethos, with users framing it as a “vital addition” for tool/server integration workflows.
- Clef: Open Weights decision model by Cloudflare (opens in new tab) (Activity: 619): Cloudflare announced Clef, an open-weights “decision model” intended for local/self-hosted use. A top commenter notes that Clef was post-trained from Qwen3.8-27B and that clef-flash was also released, post-trained from Qwen3.5-9B. The main substantive reaction was positive: commenters see Clef as filling a gap in the local-model ecosystem and are eager to benchmark it themselves.
-
Commenters noted that Cloudflare Clef is post-trained from Qwen3.8-27B, with a smaller clef-flash variant post-trained from Qwen3.5-9B, framing it as a potentially important open-weights “decision model” for local inference use cases.
- A technical concern raised was how Clef’s quality holds up after quantization, especially below Q8, since local deployment will likely depend on lower-bit quantized variants and decision-model behavior may degrade nonlinearly under aggressive compression.
- The benchmark discussion focused on comparisons against models such as Laya, Kev 9B, and DiffusionGemma Jev, but one commenter criticized the eval set as too weak and argued Clef should be compared against the leading models on jevbench rather than weaker open Jev baselines.
-
Commenters noted that Cloudflare Clef is post-trained from Qwen3.8-27B, with a smaller clef-flash variant post-trained from Qwen3.5-9B, framing it as a potentially important open-weights “decision model” for local inference use cases.
- New in llama.cpp: Decision Models (opens in new tab) (Activity: 574): The post announces Decision Models support in
llama.cpp: local “Jev/Jeff-like” models intended to act more like controllers/classifiers—selecting among actions, continuations, or behavioral choices—rather than purely free-form generators. No benchmark numbers or low-level implementation details were discussed in the provided comments; the main concrete use case raised was steering local roleplay models to avoid characters “going off the rails mid scene.” Commenters were skeptical about Jev as a defensible product/category, arguing the idea had “no moat” and was rapidly cloned into many Jev-like models. Others said they still do not know what these models are practically useful for, aside from possible agent/roleplay control.-
A commenter frames decision models as essentially a constrained classification loop: provide a JSON schema containing allowed classes plus structured/unstructured input, then have the model select exactly one category per item. The technical question is whether llama.cpp’s new support adds meaningful inference-time behavior beyond ordinary prompt-constrained JSON classification, or primarily standardizes the workflow for local models.
- One practical use case raised is applying decision models to local roleplay agents to reduce derailment during long scenes—i.e., using an auxiliary model or decision step to enforce state/intent constraints before generation. The thread does not report benchmarks or implementation results, but highlights a potential control-layer pattern for character consistency and scene-state management.
-
A commenter frames decision models as essentially a constrained classification loop: provide a JSON schema containing allowed classes plus structured/unstructured input, then have the model select exactly one category per item. The technical question is whether llama.cpp’s new support adds meaningful inference-time behavior beyond ordinary prompt-constrained JSON classification, or primarily standardizes the workflow for local models.
Less Technical AI Subreddit Recap
/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo
1. Gemini 4 Argon Access Backlash
- Just canceled my Google One AI plan. (opens in new tab) (Activity: 1838): The OP claims Google’s paid Google One AI Pro tier no longer provides access to frontier Gemini models: after an alleged Gemini 4 Argon announcement, access is described as limited to enterprise “Fairwind” partners, paid API users, and a forthcoming Google AI Ultra tier, while Pro users remain on Gemini 3.8 Flash. They argue this is a regression from the Gemini 2.5 Pro era—where higher-end reasoning models and generous limits were available more broadly—and contrast it with Anthropic/OpenAI subscriptions allegedly offering frontier models to standard paid users; no concrete benchmark numbers are provided beyond claims of “impressive benchmark charts” and a
1Mtoken output ceiling. Top comments mostly dismiss the complaint: one user says they subscribe primarily for Google storage and treat AI as a bonus, while others question the post’s authenticity, alleging it was Gemini-written or bot/shill activity from a new account.-
One commenter argued that the
$20/monthAI subscription tiers from Google/Anthropic/OpenAI function more like constrained trials than production-grade access, implying practical limits on sustained workloads despite “Pro” branding. They also suggested Google may be subsidizing or losing money on Google One AI Pro subscriptions given the underlying inference costs.- A rollout clarification noted that Gemini Ultra appears to be receiving access first, but Google has not explicitly ruled out Pro-tier access to features like Astra. The commenter framed this as a typical staged software rollout rather than definitive permanent tier exclusion.
-
One commenter argued that the
- Why publicly announce a model that the public can’t use yet?? (opens in new tab) (Activity: 1624): The image is a screenshot of a purported Google/Gemini announcement for “Gemini 4 Argon”, claiming frontier performance in software engineering, knowledge work, and cybersecurity defense, plus an extremely large
1M token output limit: image (opens in new tab). The post’s technical significance is mainly about model-release communication, not evaluation: the title questions why Google would publicly announce a model before it is accessible to users, and the image itself provides no benchmarks, API details, pricing, or availability timeline. Commenters compare this to prior “announced but unavailable” model rollouts, including Anthropic’s “mythos” and Google’s alleged “3.5 pro” handling. The dominant view is skeptical: users may be frustrated, and one commenter speculates the announcement is aimed “purely for investors.”
2. Claude Opus 5.5 Regression Reports
- Opus 5.5 nerfing - how to measure, how to spot, how to sue (opens in new tab) (Activity: 2722): Poster alleges Anthropic Opus 5.5 showed a sharp post-launch regression after
5–6days on complex C++/3D/physics/Blender MCP workloads, citing abnormal phrasing and lower code/output quality, and recommends preserving exact launch-day prompts/outputs plus latency measurements to detect potential changes such as quantization, routing, or serving optimizations under load. They frame this as a potential EU consumer-law issue under the Digital Content Directive 2019/770 (opens in new tab), specifically conformity expectations in Arts.7–8and modification/withdrawal notice obligations in Art.19, arguing launch benchmarks and “most capable model” marketing may set enforceable expectations. Comments broadly agree that closed-model providers can silently degrade or reroute models and that independent auditing is needed, but no commenter provides reproducible benchmarks or direct evidence. One commenter reports similar perceived quality drops in Higgsfield outputs, describing wasted credits after initially strong generations.-
Commenters raised the core measurement problem with alleged closed-model degradation: because Anthropic’s hosted model weights, prompts, routing, and serving configs are opaque, users argue it is difficult to prove a regression or “nerf” without independent auditing, fixed benchmark prompts, repeated sampling, and historical baselines. One user specifically asked for an “effective test or reliable nerf tracker site,” highlighting demand for third-party longitudinal evals rather than anecdotal comparisons.
- Several users reported anecdotal regressions in Opus 5.5 behavior across applied workflows: one claimed it now needed help from Gemini 3.8 Flash to catch coding bugs, while another said Higgsfield design/render outputs declined after initially strong results, wasting credits. These reports are not controlled benchmarks, but they point to the kinds of tasks users want tracked: bug-finding accuracy, design/render prompt fidelity, and day-over-day output consistency.
-
Commenters raised the core measurement problem with alleged closed-model degradation: because Anthropic’s hosted model weights, prompts, routing, and serving configs are opaque, users argue it is difficult to prove a regression or “nerf” without independent auditing, fixed benchmark prompts, repeated sampling, and historical baselines. One user specifically asked for an “effective test or reliable nerf tracker site,” highlighting demand for third-party longitudinal evals rather than anecdotal comparisons.
- Mmmkay. I didn’t believe others at first, but something is suddenly off with Opus 5.5 (opens in new tab) (Activity: 2045): A Claude Code Enterprise PAYG user reports a sharp perceived regression in Claude Opus 5.5 Med behavior after a monthly limit reset: from architecture-first, DRY/SOLID, token-efficient implementation to verbose preambles, duplicated code, “slopcode,” and token burn resembling prior Opus 5 behavior. They claim usage jumped from roughly
70%to90%in about an hour, versus no spend-limit increase requests during the previous week of heavy~12h/dayO5.5 use, and offer daily cost/token data for comparison. A commenter cites external sentiment tracking showing Opus 5.5 Reddit sentiment dropping from71–73/100on Sep 25–28 to58yesterday and55today on modelsentiment.com (opens in new tab), while noting it measures opinion rather than backend model changes. Top comments speculate Anthropic may have reduced compute, silently changed routing, or altered token accounting after launch hype, but no direct evidence is provided. The main debate is trust/reliability: users want stable model behavior and transparent deployment/versioning rather than perceived post-release regressions.-
A commenter tracking Reddit sentiment reports a sharp drop for Claude Opus 5.5, with scores allegedly stable at
71–73/100from Sep 25–28 before falling to58yesterday and55today on modelsentiment.com (opens in new tab). They note this measures user opinion rather than model behavior, so it cannot confirm a backend change, but it may indicate a sudden perceived quality regression.- Multiple users describe a suspected capability regression in Opus 5.5, especially around instruction-following and multi-part prompt adherence: one says the model now “mentions 3 things and only acknowledges 2,” and even recognizes the omission when challenged. Another user says they reverted to “xhigh effort” mode for all tasks, implying lower default reliability or reduced reasoning/compliance under normal settings.
- One technical hypothesis raised is that Anthropic may have temporarily allocated more compute during launch/benchmarking and later reduced inference resources or altered token accounting, leading to perceived quality degradation. This is speculative and unverified, but the complaint centers on reproducibility and reliability: users want model behavior to remain stable after release rather than changing silently under the same product name.
-
A commenter tracking Reddit sentiment reports a sharp drop for Claude Opus 5.5, with scores allegedly stable at
3. AI Video Models and Motion Control
- Orbiting Lora + first and last frame in MiniMax gives fantastic results (opens in new tab) (Activity: 2263): A user shared a MiniMax-H3 LoRA for generating locked-subject
360°orbit shots from first/last-frame conditioning:pablodawson/MiniMax-H3-360-Orbit-LoRA. The prompt explicitly constrains the scene to a frozen instant—no object/pose deformation, no drifting, no continued action—so that camera parallax is the only motion source, targeting cleaner pseudo-volumetric outputs suitable for downstream reconstruction workflows. A linked Reddit demo video was mentioned, but the video URL could not be inspected due to Reddit returning403 Forbidden. Commenters framed the LoRA as especially useful for creating 3D assets: one suggested feeding the generated orbit clip into Opus to extract snapshots for 3D model generation, claiming it improves style preservation. Another commenter extrapolated that this kind of orbit-consistent video generation brings consumer volumetric/VR viewing of existing films closer.-
One commenter describes a workflow where an orbiting/generated 3D video snippet is fed into Opus and instructed to extract snapshots for 3D model creation. They report that this improves style capture substantially versus prompting the model without the video reference, suggesting the orbit video acts as a strong multi-view conditioning source.
- A technical artifact noted in the output is inconsistent motion segmentation: humans remain effectively frozen while secondary elements such as the car, hair, and background explosion continue moving. This points to MiniMax preserving the subject pose from the first/last-frame constraints while still synthesizing environmental dynamics, which can create partial-animation mismatches.
-
One commenter describes a workflow where an orbiting/generated 3D video snippet is fed into Opus and instructed to extract snapshots for 3D model creation. They report that this improves style capture substantially versus prompting the model without the video reference, suggesting the orbit video acts as a strong multi-view conditioning source.
- Griffin, the first Human Interaction Model to pass video Turing Test it’s already #1 on NVIDIA’s benchmark for full-duplex AI video - 44% of people thought it was a real person while other systems are at ~3% (opens in new tab) (Activity: 2018): A Reddit post claims Griffin, described as a “Human Interaction Model,” is the first system to pass a video Turing Test and ranks
#1on NVIDIA’s benchmark for full-duplex AI video, with44%of participants judging it as a real person versus roughly~3%for other systems. The linked Reddit video could not be independently accessed due to a 403 Forbidden response, so the benchmark details, methodology, and model architecture are not verifiable from the provided source. Comments were mostly non-technical: one user joked about the human/AI reveal being reversed, while another argued the technology is unnecessary and likely to be used in predatory applications.-
A commenter emphasized that a
44%human-identification rate is technically significant because humans are usually highly sensitive to subtle facial, timing, and behavioral anomalies—the basis of the uncanny valley problem in CGI/animatronics. They argued this suggests Griffin is substantially beyond prior “fake human” systems, especially compared with the post’s claim that other systems score around~3%on the same video Turing-style benchmark.- One technically relevant real-world abuse case raised was AI-generated job applicants: synthetic candidates allegedly apply, conduct video interviews, get hired, and then either gain internal platform access or steal shipped work equipment. The commenter noted that large companies with weak scrutiny or limited background checks may fail to detect these AI-mediated interviews, implying full-duplex video agents could materially worsen identity-verification and hiring-security risks.
-
A commenter emphasized that a
- Pi 1.0 added native MCP support to Codemode, deferred tool loading, Anthropic cache warming, mid-conversation system messages, and support for non-LLM/image models and virtual-model extensions. Claude Code mods expose middleware-like hooks; its “You should know” plugin runs a side agent that flags important output the user might miss.
-
T3 Code’s rewrite added cross-provider
delegate_task, thread forking, mid-thread model switching, subagent lineage views, scheduled tasks, Pi support, and an ACP registry. Cursor Rollouts can identify the PR behind a regression, open an issue, and offer a one-click cloud-agent fix. - Agent performance can vary substantially with the harness: identical weights scored 62% in one harness and 33% in another. Multi-harness training raised LFM2.5-2.6B from 42% to 54% across four harnesses and cut tool calls by 31%; the trainer, data, and seven trained models were released openly. Meta Superintelligence Labs also reported that a dedicated controller raised GPT-5.5’s ProgramBench score from 63.7% to 71.5% with the same workers and budget, versus 58.0% for Codex.
-
A user tested Qwen3.8-27B IQ4_XS via
llama.cppon one 24 GB RTX 4090 at a 196,608-token context, reporting 12/12 on a code-review task and 40/43 hidden tests on one DeepSWE task—but the binary pass score was 0, and the author cautioned this did not establish broad parity across the 113-task benchmark.