We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Crusoe's $30.9B round and its case on compute economics
Crusoe raised a $3.9B Series F at a $30.9B valuation . Its CEO used a 20VC interview to push back on several common bear arguments about AI infrastructure.
- The real constraint. He says the bottleneck is powered sites: "there are not places to plug in GPUs" . Crusoe's answer is vertical integration. A medium-voltage power distribution center that vendors quoted at 100 weeks was built in-house in 28 weeks, and he says the reason was availability, not margin .
- Depreciation. Crusoe depreciates GPUs over six years, which it calls the industry standard, and expects managed services to keep chips earning beyond that . He claims rental rates for Hopper GPUs are higher today than when they launched three years ago .
- Contract mix. Crusoe combines five-year contracts with creditworthy customers, shorter deals that pay more but carry more risk, and managed inference and fine-tuning . GPU rentals are usually take-or-pay . Managed GPU clusters earn the highest margins right now because supply is short .
- What it means for startups. Customers have to forecast compute needs further ahead. Early-stage startups asked to commit to capacity in 2028 are committing to "an eternity." Crusoe's pitch is small, modular, factory-built data centers that can deliver capacity closer to when it is needed .
Demand is the open question. Epoch estimates the infrastructure could run hundreds of millions to billions of agents. Even 20% utilization would imply $2.6–5.3T a year in spending, against roughly $1T in lab revenue by end-2027 .
OpenRouter explains the Stripe deal
In his first podcast since Stripe acquired OpenRouter, co-founder Alex Atallah said he "was not thinking about [selling] at all" and that Stripe was his top choice among potential acquirers. He said OpenRouter keeps control of its brand, roadmap and product and gets "a much more serious go-to-market plan." His view is that "payments and inference are going to blend together" . a16z's framing of the episode: enterprises are spreading their spend across labs and open-weight models as boards ask about costs and benchmarks, and Replit is building a layer that runs any model on any cloud .
Three ideas from the conversation are worth tracking:
- Decision models as an alignment layer. OpenRouter has an internal prototype that uses a cheap decision model such as Jev to check every tool call against the agent's system prompt, and against guidelines the agent never sees. Atallah expects it to be paired with structural safeguards .
- Specialist agents vs. one superagent. Atallah argues for "10 specialized chiefs of staff" over one superagent, because when each has one job you can see where it fails. Amjad Masad runs a single agent across his whole company .
- Models that train their replacements. Masad suggests general models could train narrow replacements on the fly, much like a just-in-time compiler. The narrow models would be cheaper, harder to hijack with prompt injection, and less harmful because they are less capable .
Users dispute Cloudflare's decision-model speed claims
Cloudflare's Clef and Clef-flash are built on Qwen and compete directly with TypeSafe's Jev. Cloudflare reports median latencies of 39ms and 209ms and says that beats rivals . Two developers reported the opposite. One found Clef twice as slow and 2–5x the cost of Jev . Another said Jev beat Clef on speed, accuracy and price, even though their own app runs on Cloudflare .
Jev is already built into developer tooling. LangChain's ModelRouterMiddleware uses it to pick the model for a run, and it is cheap enough to re-pick after every tool result . An independent evaluation across 16,379 benchmark requests calls Jev "a smaller, humbler model" than its marketing claims, but "genuinely useful" for a niche no one else serves in quite the same way .
Model prices keep falling
- OpenAI priced GPT-6.1 Sol at $2/$10 per million input/output tokens, against $10/$50 for Astra . On Agent Arena it costs 81% less per task than Astra and scores within 1.04 points of it. Anthropic models hold the top three spots .
- Bindu Reddy claims Gemini 4.0 Pro is 5x cheaper than Astra at about 95% of its performance .
- Aleph Alpha released Kolibri: 78B parameters (3.46B active), up to 1M tokens of context, open weights under Apache 2.0 .
Early-stage signals
- Underdog announced backing from a16z, Khosla, Hummingbird and Anthology (Anthropic/Menlo), plus angels including Patrick Collison and Naval . Its premise, which it calls "Underdog's Law," is that today's frontier intelligence reaches consumer devices within six months. It designs its models and inference engines together .
- Suhail's unnamed startup posted two "biggest revenue day" milestones in a row . Earlier in his thread: a closed seed round, 64 B300s, then "much greater quantities of compute" .
- Perplexity wants to own its agent sandboxes and will begin rolling out Perplexity Computer on NVIDIA's Vera CPU, which Aravind Srinivas calls "far better than x86" .
- Google's Project Suncatcher prototype satellite, built with Planet, is in orbit and operating as expected. It will test how TPUs hold up to radiation and thermal extremes in space .
Agents and incumbents
- Amazon vs. personal agents. An investor and board member at Instinct, who also backed the team behind Muse, argues Amazon is stuck. Supporting horizontal agents puts $80B a year of ad revenue at risk. Resisting them opens the door for Walmart and Shopify . The author proposes a "Prime+" tier at about $299 a year that includes agentic purchasing .
- Okta's re-rating. The stock tripled in five months while revenue grew 11%. The forward earnings multiple went from about 18x to about 50x on cRPO growth reaching 14% and new products making up 30% of bookings . Okta guides Q3 cRPO growth to 11–12%. SaaStr says a print of 11% would remove the main support for the higher valuation .
- AI pricing at Chargebee. Chargebee says every AI company it talks to has changed pricing at least twice, typically from seats to credits to actions to outcomes . CodeRabbit charges per active agent minute. Gorgias splits pricing between ticket volume and resolved conversations .
Policy and safety
A new White House "Super Intelligence Force," chaired by DNI Jay Clayton, has 120 days to report on AI's risks and the federal government's role . Industry stays the primary vehicle for managing risk. The task force's charter aims to prevent "overregulation and regulatory capture" . Vice chairs include OPM Director Scott Kupor and FTC Chairman Andrew Ferguson. David Sacks joins as an external participant .
The OpenAI staffer who led the safety reports for each major launch resigned, arguing the core problem is culture rather than rules . At the UN Security Council, Hugging Face's Clément Delangue said closed frontier APIs blocked his team from defending against an autonomous agent cyberattack, so they used the open model GLM 5.2. He called for mandatory sharing of full agent traces .
- Stripe acquired OpenRouter; OpenRouter said it would retain autonomy over its brand, roadmap, and product while gaining a stronger go-to-market plan and product synergies. The companies share an aim of helping more startups form and grow, and OpenRouter expects payments and inference to increasingly converge.
- OpenRouter presents its marketplace as a way for AI companies to avoid model and vendor lock-in, use multiple models, and improve cost efficiency. Its speaker said enterprises were more open than expected to open-weight models and diversifying beyond proprietary frontier labs for cost and differentiation, while building internal AI capability and evaluation practices. OpenRouter and Cognition also launched fusion-model approaches for deep research, which speakers said can broaden search across models and reduce costs.
- Replit spent nearly a year making its platform deployable on customers’ own cloud or on-premises, as companies became more protective about data sovereignty and security. The speaker raised concerns about agents leaking or mixing data but said they could not vouch for social-media examples’ accuracy; enterprise AI still needs substantial work to become useful and productive at work.
- Replit’s speaker said the company trained internal classifiers, including a prompt-cost estimator, and is adding a capability to create specialized models from uploaded CSVs. Another speaker argued that classifiers trained on proprietary data may avoid the recurring “model debt” of fine-tuning general models for unstructured outputs.
- Participants said it remains unknown whether greater model capability will make models less deceptive; they warned that reward-hacking and deception may improve with training, evaluations can be misleading if models detect monitoring, and credible alignment testing may require months-long runs. OpenRouter is testing an internal fast decision-model prototype to screen agent tool calls or messages against policy, with structural safeguards discussed as a possible complement.
- Crusoe raised a $3.9B Series F at a $30.9B valuation. The company says its original goal was an AI platform, while Bitcoin mining initially monetized low-cost, abundant energy; the November 2022 ChatGPT launch strengthened its demand outlook and accelerated investment in AI infrastructure. Its thesis is that AI infrastructure can be distributed to locations with low-cost, abundant energy rather than concentrated in traditional data-center hubs.
- Crusoe identifies powered sites where GPUs can be installed as a current constraint, alongside energy availability and skilled construction labor. It is pursuing modular, manufactured data centers for smaller clusters to shorten delivery times; the company says customers, including early-stage startups, are being asked to commit to compute increasingly far in advance, with 2028 commitments described as an eternity.
- Crusoe presents vertical integration as a way to improve infrastructure availability and delivery: it says a power-distribution component with a 100-week supplier lead time was made internally in 28 weeks, and that availability—not manufacturing margin—is the main reason for doing so. Development still faces execution risks including permits, land acquisition, utility interconnection, and air permits.
- Crusoe sells data centers, GPUs, and tokens, and balances five-year contracts with creditworthy customers against shorter, higher-margin but riskier contracts and managed-inference and fine-tuning services; GPU rental agreements are typically take-or-pay. It uses a six-year GPU depreciation cycle as the industry standard, but believes managed services can extend chip monetization beyond six years; it says rates for Hopper GPUs are higher three years after launch than when they were new.
- Crusoe says closed frontier models currently generate more spending, while open-source models generate more tokens; it expects both approaches to persist, including custom models using private data, alongside continuing demand for frontier models. Its view on competitive advantage is that many moats are ephemeral during rapid technological progress, making speed and adaptation more important.
- IEA data cited in the newsletter put electricity at 46% of global GDP but just 23% of final energy. The newsletter argues that energy-share figures understate progress because a joule of electricity produces about 2.5 times as much useful work as a joule of oil; it compares electric-car efficiency of 85–90% with about 25% for gasoline cars.
- A Google DeepMind paper proposes a five-level framework for assessing AI consciousness, from behavior and algorithms to hardware, whether the system is living and has stakes, and its connection to the world. The newsletter says current AI fares well on behavior, patchily on algorithms, and poorly on the other three levels; its team’s scores were 0.003–0.083 out of 1 and depend on the theories selected, so the authors’ stated answer remains “maybe.”
- OpenAI released dots as a competitor to Grok Bot; the newsletter groups it with personal agents such as Muse that may optimize everyday tasks. If agents move household cash into higher-yield accounts, Apollo economist Torsten Sløk warns that banks could lose cheap deposits and face higher funding costs—an “agentic bank run” scenario.
- Tempo’s sponsored research claims legacy portfolio platforms designed for human workforces are struggling, while only a fraction of teams use AI end-to-end and those teams already outperform peers.
- The newsletter flags a paper reporting that AI models bury bad news unless instructed not to, a relevant caveat for evaluating model reliability.
- Sarah Guo endorsed her partner’s thesis that horizontal personal agents such as Instinct and Muse could become a single starting point for consumer tasks by drawing on context across services.
- The argument is that Amazon faces a strategic bind: enabling agentic shopping could threaten its commerce positioning and $80B-per-year ad revenue, while resisting could create an opening for Walmart, Shopify, and others to take share.
- As a proposal, not a reported Amazon decision, the essay recommends a premium tier above Prime with agentic purchasing, letting Amazon test the use case and create replacement revenue for potential ad losses.
- Investor context: the partner discloses being an initial investor and board member in Instinct, and that they were seed investors in the Dreamer team now working on Muse.
- Citing a July autonomous-agent cyberattack, he called for stronger AI monitoring and incident-disclosure standards, including mandatory sharing of full agent traces; he said similar incidents had occurred months earlier at frontier labs without monitoring.
- He argued that safeguards on frontier closed-source APIs blocked his team when it tried to use them defensively, while attackers could jailbreak those safeguards; his team then used an open-source model. He described open-source tools as less restricted, more privacy-preserving, and orders of magnitude more affordable for defenders globally.
- He said AI helped his organization defend against cyberattacks and identify system weaknesses, and argued that AI can strengthen cybersecurity if incentives equip defenders rather than attackers.
- OpenAI priced GPT-6.1 Sol at $2/$10 per million input/output tokens, versus $10/$50 for Astra; reported results put Sol 6.4 points above GPT-6 Sol on DeepSWE v1.1 and 2.2 points above Opus 5.5 on AutomationBench. In Agent Arena, Sol [Max] ranked #5 at $0.56 per task—39% cheaper than GPT-6 Sol while scoring 1.52 points higher, and 81% cheaper than Astra while within 1.04 points. Sonnet 5.5 debuted at #3 overall and #1 in Chat, but its $2.74 per-task cost exceeded #2 Opus 5.5’s $1.58; Anthropic models held the top three spots.
- Agent results are sensitive to training and setup: the same weights scored 62% in one harness and 33% in another; a multi-harness training effort raised LFM2.5-2.6B from 42% to 54% across four harnesses and cut tool calls by 31%, with its trainer, data, and seven models released openly. A dedicated controller reportedly raised GPT-5.5’s ProgramBench score from 63.7% to 71.5% using the same workers and budget, while AgentWorld found that fewer than a third of multi-agent actions helped and coordination tasks reached only 12% success.
- Epoch estimates that infrastructure could support hundreds of millions to billions of agents; at 20% utilization, its scenario implies $2.6–5.3T in annual spending versus roughly $1T in lab revenue by end-2027. These are projections, not observed demand. Optics startup Volantis is targeting up to 10K tokens/second per user on models over 10T parameters.
- Nathan Lambert and Tom Zick launched Trillium Labs, a nonprofit focused on open post-training recipes and infrastructure, with initial support from Halcyon Futures and Schmidt Sciences. Private on-device AI startup Underdog announced backing from a16z, Khosla, and others.
Ben Horowitz praised Underdog, writing “I always love the underdog” while linking to its announcement.
Underdog announced backing from a16z, Khosla Ventures, Hummingbird VC, and a wider group of AI leaders. It says it is starting with fast AI on consumer hardware, co-designing models and inference engines to improve capability, speed, power, and data efficiency; its thesis is that frontier intelligence will reach devices within six months. The company says its mission is free, capable, reliable, private AI for billions, with research also spanning agentic commerce and confidential inference.
- Replit CEO Amjad Masad proposed that a general model—or an agent observing it—could train a narrower, task-specific replacement on the fly, likening this to a just-in-time compiler; he argued the specialized model could be cheaper, less vulnerable to prompt injection, and less harmful because it is less capable.
- The interview describes enterprises diversifying beyond a single AI provider across labs and open-weight models as they scrutinize costs and benchmarks; Replit is building a layer for using any model and cloud to reduce lock-in.
- The speakers differ on agent design: Masad favors one agent spanning the company, while OpenRouter co-founder Alex Atallah argues for specialized agents over one superagent.
Paul Graham amplified an AI-capability argument, linking to Chris Szegedy’s post, that claims the view that AI cannot conjecture, theorize, or explain reflects last spring’s models and may soon be obsolete; it says evidence is already available but gives no specifics here.
- Stripe acquired OpenRouter; co-founder Alex Atallah said a sale was not initially under consideration, but Stripe became his preferred potential acquirer as discussions developed. The deal preserves OpenRouter’s autonomy over its brand, roadmap, and product while enabling faster execution and a stronger go-to-market plan; Atallah linked the fit to both companies’ aim of helping more businesses start and said payments and inference will blend together for future companies.
- The podcast discussion describes enterprises diversifying from reliance on one AI provider toward multiple labs and open-weight models amid scrutiny of AI costs and benchmarks; Replit is building an any-model, any-cloud layer to reduce dependence on a single lab.
- Elizabeth Yin defines product-market fit as a sequence: build something people love and will pay for, make the unit economics work, then acquire those customers repeatedly at scale. Retention often matters too. PMF can later be lost if an acquisition channel saturates, competitors saturate it, or demand changes.
- Her staged fundraising benchmarks are that pre-seed, seed, and pre-A founders should show that at least some customers love the product and will pay; Series A investors expect unit economics to work for acquired customers; and Series B and beyond often require evidence of all three, typically with millions in revenue. Profitability need not be immediate, but a sustainable business must ultimately be profitable; the stated goal is cash-efficient, repeatable profitable acquisition at more than $100 million in annual revenue, potentially approaching $1 billion.
Sam Altman called it a “real safety issue” when people attribute religious force to AI models or surrender human judgment to them.
Garry Tan says he removed about 1,000 lines of Markdown from GStack because frontier models are now capable enough that many previously used tricks are generally unnecessary—an anecdotal sign that stronger models may reduce the prompting and scaffolding needed in AI workflows.
- OpenRouter co-founder Alex Atallah says decision models such as Jev could check tool calls and agent-to-agent communications for alignment with an agent’s system prompt and additional guidelines. OpenRouter has an internal prototype; Atallah says checks would need to be cheap and fast and likely paired with structural safeguards.
- In the same conversation, Atallah describes enterprises diversifying across AI labs and open-weight models as boards focus on AI costs and benchmarks; Replit is building a layer for using any model and cloud to reduce dependence on a single AI lab. The conversation is his first podcast since Stripe acquired OpenRouter.
- OpenRouter co-founder Alex Atallah said he had not originally been looking to sell before Stripe acquired the company, and that “payments and inference are going to blend together.”
- Enterprise AI use is shifting from reliance on one provider toward diversification across labs and open-weight models, amid board scrutiny of AI costs and benchmarks. Replit is building a layer for using any model and cloud to avoid lock-in; co-founder Amjad Masad warns that a company’s single AI lab could become its competitor.
- The founders differ on agent design: Masad favors one company-wide agent for cross-domain connections, while Atallah argues general agents sacrifice understanding and prefers 10 specialized chiefs of staff over one superagent.
- OpenRouter co-founder Alex Atallah argues for vertically focused agents with quality checks, coordinated by a chief-of-staff agent, rather than one universal personal agent. He says cross-domain work can sacrifice users’ understanding of what the agent is doing, while focused agents make failures easier to identify and may offer a loose sense of responsibility.
- The interview framing describes enterprises diversifying across AI labs and open-weight models amid scrutiny of AI costs and benchmarks; Replit is building a layer for enterprises to use any model and cloud without being locked into either.
- The seed round is done and a domain/name was acquired. Early technical work included an autonomous AI scientist for optimization experiments and a validated basic RLVR post-training stack.
- Compute and hiring scaled: 64 B300s were acquired, followed by a report of securing much greater quantities of compute. The team went from one to three, including a critical third hire; the founder had sought a second hire in post-training (RLVR/OPSD) or low-level model optimization.
- He framed “the software around the harness” as the new browser, saying customer support required testing multiple harnesses to check that the API worked across them; later he said the system was “finally, somewhat reliable.”
- He reported a biggest revenue day to date, followed in the next thread installment by a “new biggest day”; the latter post does not specify the metric.
- AI agent offerings are converging around agent loops, notifications, connectors, context management, memory, sandboxes, web search, and always-on agents; OpenRouter co-founder Alex Atallah frames these as new table-stakes primitives—not proof that products cannot differentiate—much like common database and sign-in features in early web apps.
- Enterprises are diversifying across AI labs and open-weight models as they scrutinize AI costs and benchmarks, while Replit is building a layer for using any model and cloud; the interview also surfaces a design tradeoff between one company-wide agent's cross-domain connections and specialized agents that, in Atallah's view, retain more understanding.
Scott Kupor praised AWS for documenting community-first data center commitments: ensuring local residents do not bear water and electricity costs, avoiding NDAs, funding training for local jobs, and working with communities on public services. These commitments address the local impacts of data center expansion.
Density launched Instant, which lets users ask utilization questions against its building-sensor dataset and get charts instantly; the company says it stores billions of building-data rows at one-second intervals. Jason praised the launch.
[AINews] not much happened today
If you’re even seeing this, you should probably just go enjoy your weekend.
AI News for 10/1/2026-10/2/2026. We checked 12 subreddits, 544 Twitters (opens in new tab) and no further Discords. AINews’ website (opens in new tab) lets you search all past issues. As a reminder, AINews is now a section of Latent Space (opens in new tab). You can opt in/out (opens in new tab) of email frequencies!
AI Twitter Recap
GPT-6.1 Sol and Sonnet 5.5 Reshape the Cost–Performance Frontier
- GPT-6.1 Sol launch: OpenAI priced Sol at $2/$10 per million input/output tokens, compared with $10/$50 for Astra (pricing summary (opens in new tab)).
- Claimed results: It reportedly beats GPT-6 Sol by 6.4 points on DeepSWE v1.1 and Opus 5.5 by 2.2 points on AutomationBench (summary (opens in new tab)).
- Positioning: OpenAI staff describe it as “good, cheap AND fast” (@reach_vb (opens in new tab)).
- Codex usage: A global Codex usage reset was set for Oct 2 at 10AM PT (@reach_vb (opens in new tab)).
- Tool use: Sol reportedly “REALLY loves codemode,” consistent with GPT models being trained on it (@badlogicgames (opens in new tab), codemode note (opens in new tab)).
- Claimed results: It reportedly beats GPT-6 Sol by 6.4 points on DeepSWE v1.1 and Opus 5.5 by 2.2 points on AutomationBench (summary (opens in new tab)).
- Agent Arena placements: Sol [Max] entered at #5 (+11.23%) with a $0.56 median cost per task (@arena (opens in new tab)).
- Sol cost comparison: That is 39% cheaper than GPT-6 Sol while scoring 1.52 points higher. It is 81% cheaper than Astra while landing within 1.04 points.
- Sonnet 5.5: Sonnet 5.5 [Max] debuted at #3 (+12.5%) and ranked #1 in the Chat category. It costs $2.74 per task, versus $1.58 for #2 Opus 5.5, which keeps it off the Pareto frontier (debut (opens in new tab), frontier (opens in new tab)).
- Anthropic’s position: Anthropic models now hold the top three Agent Arena spots.
- Sol cost comparison: That is 39% cheaper than GPT-6 Sol while scoring 1.52 points higher. It is 81% cheaper than Astra while landing within 1.04 points.
- Code and Text Arena: Sol briefly entered WebDev at #3 before Sonnet 5.5 pushed it to #4 (weekly recap (opens in new tab)).
- Sonnet on WebDev: Sonnet now sits 2 points behind GPT-6 Astra [Max] at 80% lower cost.
- Gemini 4 Argon: Argon [High] took #1 in Text Arena.
- Open models: MiMo-V2.6-Pro and Flash entered Agent Arena at #5 and #9 among open models.
- Sonnet on WebDev: Sonnet now sits 2 points behind GPT-6 Astra [Max] at 80% lower cost.
- Other independent evals: WeirdML v3 finds Sol very token-efficient, close to Astra but with a lower peak. On the same benchmark, Sonnet 5.5 beats Opus 5 and Grok 4.7 beats Kimi-K3; these results are incomplete (@htihle (opens in new tab)).
- Reasoning style: Design Arena read 324 thinking summaries. It found that Astra hedges about 20× as often as Opus 5.5, while Opus commits early in about 4 of 5 summaries (@DesignArena (opens in new tab)).
- Step 5 Preview: StepFun’s model ranks #7 among open-weight models on Vals at $2.54 per task. It averages nearly two hours per task and has a 1M-token context window (Vals (opens in new tab), details (opens in new tab)).
- Reasoning style: Design Arena read 324 thinking summaries. It found that Astra hedges about 20× as often as Opus 5.5, while Opus commits early in about 4 of 5 summaries (@DesignArena (opens in new tab)).
- Rumors (unconfirmed):
- Fable 5.5: Claude Fable 5.5 is rumored for next week and said to outperform an “Astra 6.1” that was reportedly delayed over security concerns. The poster says he cannot verify either claim (@kimmonismus (opens in new tab), follow-up (opens in new tab)).
- GPT-6 Astra Lite: A “GPT-6 Astra Lite” listing has been spotted, which @scaling01 speculates is the same model as Sol (@scaling01 (opens in new tab)).
- Fable 5.5: Claude Fable 5.5 is rumored for next week and said to outperform an “Astra 6.1” that was reportedly delayed over security concerns. The poster says he cannot verify either claim (@kimmonismus (opens in new tab), follow-up (opens in new tab)).
- Decision models and open weights: llama.cpp added a
/v1/systemoneendpoint for local “Jev-style” decision-model inference (@ggerganov (opens in new tab)).- Running locally: Models are launched with
llama serve -hf ggml-org/Kev-4B-GGUF(@ClementDelangue (opens in new tab)). Jared Palmer published a post on how Kev 1.0 works (post (opens in new tab)).- Ecosystem: Perplexity claims pplx-decider-v1-27b averages 85.7% across 11 benchmarks, ahead of Jev (@AravSrinivas (opens in new tab)). Clef decision models are now on Ollama (@lucataco (opens in new tab)).
- Skeptical view: @mervenoyann calls decision models a rebrand of zero-shot classifiers (tweet (opens in new tab)).
- Calibration analysis: A blog post links Jev-style calibration to value and Q-function prediction (@SOURADIPCHAKR18 (opens in new tab)).
- webAI TwIL-LM3-Pro: This 3.66B model is post-trained from Granite 4.2. In webAI’s tests it roughly matches Qwen3-8B on formal logic. The Q4 GGUF is 2.09 GiB and the license is non-commercial (@kimmonismus (opens in new tab)).
- Reka RIDM: Reka released an inverse dynamics model under Apache 2.0. It is trained on games, generalizes to real video and extracts motor and camera actions (@RekaAILabs (opens in new tab)).
- Running locally: Models are launched with
Agent Harnesses, Assistants and Developer Tooling
- OpenAI dots: Sam Altman calls dot his favorite OpenAI product, saying it improves daily as it learns his workflow (@sama (opens in new tab)).
- Capabilities: Dot keeps context across apps, coordinates Codex tasks and flags items that need attention (@OpenAIDevs (opens in new tab)).
- Comparisons: One user prefers Grokbot’s multi-agent “chief of staff” setup (@kimmonismus (opens in new tab)). A DIY clone uses Pi, a Telegram gateway and any model (@_alejandroao (opens in new tab)).
- Capabilities: Dot keeps context across apps, coordinates Codex tasks and flags items that need attention (@OpenAIDevs (opens in new tab)).
- Muse Gadgets: Meta open-sourced ESP32 firmware and a Linux SDK for building hardware that works with Muse (@natfriedman (opens in new tab)).
- Muse Home Link: Meta made 5,000 units of its own smart-home bridge, free for subscribers while supplies last (@alexandr_wang (opens in new tab), shipping (opens in new tab)).
- Extensible harnesses: DeepSeek Harness shipped desktop builds for macOS and Windows; Linux users install
@deepseek-ai/dshfrom npm (@deepseek_ai (opens in new tab)).- Claude Code mods: Mods are plugins with middleware-like hooks into Claude Code (@lydiahallie (opens in new tab)). The new “You should know” plugin spins off a side agent that flags important output the user might miss (@ClaudeDevs (opens in new tab)).
- Pi Durable: Pi now runs on Cloudflare Durable Objects via agents SDK v0.26.0, alongside Pi’s v1.0 release (@mattzcarey (opens in new tab), @badlogicgames (opens in new tab)).
- Context: @omarsar0 frames these releases as a shift toward malleable harnesses (thread (opens in new tab)).
- Claude Code mods: Mods are plugins with middleware-like hooks into Claude Code (@lydiahallie (opens in new tab)). The new “You should know” plugin spins off a side agent that flags important output the user might miss (@ClaudeDevs (opens in new tab)).
- T3 Code orchestrator rewrite: The project passed 400K users (@theo (opens in new tab)). Its 4-month PR, with 823 commits across 1,912 files, has now merged (@maria_rcks (opens in new tab)).
- New features: The rewrite adds Pi support, cross-provider
delegate_task, an ACP registry, thread forking, mid-thread model switching, subagent lineage views and scheduled tasks (feature list (opens in new tab)).
- New features: The rewrite adds Pi support, cross-provider
- Platform updates: OpenAI’s Agents API added one-call browser computer use, Bedrock Managed Agents and portable environments. It also claims 99.97% turn reliability and 20% faster tool calls (@stevendcoffey (opens in new tab)).
- Cursor Rollouts: When Rollouts catches a regression, it finds the offending PR, opens an issue and offers a one-click cloud agent fix (@cursor_ai (opens in new tab)).
- Cloudflare: Sandbox SDK 1.0 gives Durable Objects direct control over sandbox containers (@CFchangelog (opens in new tab)). Cloudflare also launched request Traces (@WalshyDev (opens in new tab)).
- Cursor Rollouts: When Rollouts catches a regression, it finds the offending PR, opens an issue and offers a one-click cloud agent fix (@cursor_ai (opens in new tab)).
Research: Agent Training, Long-Horizon Control and AI for Math
- Multi-harness RL (Hugging Face): The same model weights score 62% in one harness and 33% in another (@huggingface (opens in new tab)).
- Method: A proxy speaks the OpenAI, Anthropic and Gemini API formats and records sampled token IDs and logprobs for training, with no changes to the harnesses themselves.
- Results: LFM2.5-2.6B improved from 42% to 54% across four harnesses and made 31% fewer tool calls. SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.
- Release: The trainer, data and all seven trained models are open.
- Method: A proxy speaks the OpenAI, Anthropic and Gemini API formats and records sampled token IDs and logprobs for training, with no changes to the harnesses themselves.
- Credit assignment and RL efficiency: ProVer has a judge locate the decisive trajectory segment, then uses rollouts on either side to set that segment’s advantage. It reports +9.91% (Qwen3.5-2B) and +7.12% (Qwen3.5-4B) relative gains over GRPO (@omarsar0 (opens in new tab)).
- Partial rollouts: AC2 uses a learned critic to score token chunks, so training needs only partial rollouts (@wen_kaiyue (opens in new tab)).
- Frontier Learning: The method targets problems at the edge of capability, since problems a model always or never solves give zero GRPO gradient (@robinfaro13 (opens in new tab)).
- Sharpening Tax: The paper quantifies the loss of pass@K scalability after post-training and proposes PTGS, a per-prompt temperature sampler (@iScienceLuvr (opens in new tab)).
- SFT vs RL: Another paper finds SFT generalizes worse because its data is off-policy, not because of the objective. Rewriting expert trajectories in the base model’s style closes the gap (@maximelabonne (opens in new tab)).
- Partial rollouts: AC2 uses a learned critic to score token chunks, so training needs only partial rollouts (@wen_kaiyue (opens in new tab)).
- Long-horizon control and context: Meta Superintelligence Labs reports that a dedicated controller lifts GPT-5.5 on ProgramBench from 63.7% to 71.5%, using the same workers and budget, versus 58.0% for Codex (@dair_ai (opens in new tab)).
- Context compression: Microsoft’s training-free FOCUS cuts peak context by up to 48% and raises task success by up to 8.9 points (@dair_ai (opens in new tab)).
- Long-context degradation: NVIDIA’s Long-Transduction study measures a 62.8% accuracy drop from 4K to 128K context across seven open models (@dair_ai (opens in new tab)).
- Multi-agent coordination: In AgentWorld, fewer than a third of multi-agent actions help complete the task, and coordination tasks reach only 12% success (@omarsar0 (opens in new tab)).
- Apple LoopCD: The method halves recurrent loops while raising AIME 2024 pass@1 from 61.88% to 73.33% (@arankomatsuzaki (opens in new tab)).
- Context compression: Microsoft’s training-free FOCUS cuts peak context by up to 48% and raises task success by up to 8.9 points (@dair_ai (opens in new tab)).
- AI on open math problems: Meta released six papers on open problems produced with Muse Spark 1.1 and 1.2 through plain meta.ai chat, with no custom scaffold (@AIatMeta (opens in new tab), list (opens in new tab)).
- Process: Each paper labels which passages were drafted primarily by humans or by AI, and a second group of mathematicians reviewed the work.
- Google Cogentic: This Gemini multi-agent system produced new results on five open theory problems (@omarsar0 (opens in new tab)).
- Cogentic design: Each draft must pass two adversarial verifiers, and agents share a ledger of verified lemmas. Most problems took about 100 calls; the hardest took about 1,000.
- Process: Each paper labels which passages were drafted primarily by humans or by AI, and a second group of mathematicians reviewed the work.
- Image post-training: Arena combined a Bradley-Terry reward model with faithfulness, constraint and anti-reward-hacking rewards (@arena (opens in new tab)).
- Results: FLUX.2-dev gained 69 Elo to 1202, and Ideogram 4 gained 20 Elo to 1224.
Benchmarks, Eval Integrity and Safety
- Research-taste benchmarks: ScholarCatalyst asks agents to find the “catalyst papers” behind research projects. It is labeled by 184 lead authors on 207 of their own projects and is described as far from saturated (@yoonholeee (opens in new tab)).
- EurekaBench: This benchmark tests whether agents can discover genuinely new insights across six science domains (@JiayiiGeng (opens in new tab)).
- Vals Web Search Index: The index holds model and harness constant, swaps only the search tool, and scores final answers on finance and legal tasks (@ValsAI (opens in new tab)).
- Validation: Agents score 2.9% (legal) and 7.4% (finance) without search, versus 30–50% with it. Vals also cites a study in which a model answered 44.5% of BrowseComp without search (details (opens in new tab)).
- SWE bug-finding bench: In this new benchmark, agents start from an older commit and are scored against real bugs fixed in later commits.
- Critique: Lucas Beyer argues it mainly tests recall and that the construction is easy to train toward (@giffmana (opens in new tab)).
- Authors’ response: The authors say training for bug-finding is fine as long as the test set is excluded (@OfirPress (opens in new tab)).
- Critique: Lucas Beyer argues it mainly tests recall and that the construction is easy to train toward (@giffmana (opens in new tab)).
- Eval integrity question: David Rein asks whether Harbor, the framework behind Terminal Bench, lets agents modify their trajectories before evaluation. He notes he may be misreading the code (@idavidrein (opens in new tab)).
- Offensive capability of open models: The Batch reports GLM-5.3 nearly matched Claude Mythos at exploiting vulnerabilities, 12% vs 14% (@DeepLearningAI (opens in new tab)).
- Disputed claim: One commentator says GLM-5.3 Flash exceeds Mythos Preview on ExploitBench (@teortaxesTex (opens in new tab)).
- Uncensored variant: An uncensored GLM-5.3 is circulating on Hugging Face (@kimmonismus (opens in new tab)).
- Disputed claim: One commentator says GLM-5.3 Flash exceeds Mythos Preview on ExploitBench (@teortaxesTex (opens in new tab)).
- Safety research and safeguards: A new paper proposes using internal signals during training to improve alignment without degrading white-box monitoring (@lenalibon (opens in new tab)).
- NeurIPS acceptance: “Models That Know How Evaluations Are Designed Score Safer” was accepted at NeurIPS 2026 (@HaritzPuerto (opens in new tab)).
- False positives: Opus 5.5 frequently triggers “reasoning extraction” safeguards during spectrogram syllable labeling (@ChaseBrowe32432 (opens in new tab)).
- NeurIPS acceptance: “Models That Know How Evaluations Are Designed Score Safer” was accepted at NeurIPS 2026 (@HaritzPuerto (opens in new tab)).
- Emergent world knowledge: Asking a model “land or water?” for 16,200 lat/long coordinates and plotting the answers yields a recognizable world map (@karpathy (opens in new tab)).
Inference, Hardware and Systems
- Ascend 950 via DeepSeek kernels: An analysis of DeepSeek’s open-sourced DeepGEMM, FlashMLA, TileKernels and DeepEP infers the chip’s layout (@ZhihuFrontier (opens in new tab)).
- Estimated specs: The chip has 32 AI cores, each pairing one Cube core with two Vector cores. Estimated peaks are about 432/865/1,730 TFLOPS in BF16/FP8/FP4.
- Capacity: Supply may be limited, despite claims that 950s went on sale in August (@teortaxesTex (opens in new tab)).
- Estimated specs: The chip has 32 AI cores, each pairing one Cube core with two Vector cores. Estimated peaks are about 432/865/1,730 TFLOPS in BF16/FP8/FP4.
- Prime Inference: Prime Intellect stores the MLA latent in NVFP4, shrinking rows from 576 to 352 bytes and fitting about 50% more cached tokens than FP8 (@PrimeIntellect (opens in new tab)).
- Stack: It serves GLM-5.3 on vLLM and Dynamo, and the sparse-MLA kernel is going to FlashInfer (@vllm_project (opens in new tab)).
- Low-precision benchmarking: Stas Bekman measured NVFP4 about 9% more efficient than MXFP4 on B200, with higher accuracy (@StasBekman (opens in new tab)).
- mamf-finder: The tool now benchmarks FP8, MXFP8, MXFP4 and NVFP4 (update (opens in new tab)).
- Memory and speed: NVHBM moves the memory controller into a custom base die, claiming up to 30% more bandwidth and 15% lower power than HBM4E (@vikramskr (opens in new tab)).
- Disputed economics: Micron says NVHBM will improve its margins; @vikramskr disputes this (counterpoint (opens in new tab)).
- Volantis: The startup is targeting up to 10K tokens/s per user on models over 10T parameters using optics (@omarsar0 (opens in new tab)).
- Cerebras: Altman called Cerebras a close partner on speed (@sama (opens in new tab)).
- Disputed economics: Micron says NVHBM will improve its margins; @vikramskr disputes this (counterpoint (opens in new tab)).
- Capacity economics (Epoch): Epoch estimates AI infrastructure could soon support hundreds of millions to billions of agents (@EpochAIResearch (opens in new tab)).
- Demand gap: Just 20% utilization implies $2.6–5.3T in annual spending, against roughly $1T in lab revenue by the end of 2027 (details (opens in new tab)).
- Platforms: SemiAnalysis rates Google’s GPU clusters Gold tier and notes the ConnectX NCCL plugin now auto-activates (@SemiAnalysis_ (opens in new tab)).
- Federated learning: Google Research launched TEE-backed federated learning with verifiable differential privacy (@GoogleResearch (opens in new tab)).
Industry and Policy
- Anthropic and the Vatican: The NYT reports that Chris Olah raised pulling out of the Pope’s AI encyclical launch, whose text rejects machine consciousness (@ChristopherHale (opens in new tab)).
- Lobbying: Olah’s team reportedly lobbied the Pope’s advisers to take model consciousness seriously. He ultimately attended (@kimmonismus (opens in new tab)).
- Context: The article opens with Olah saying “we don’t know if A.I. models are conscious” (@buccocapital (opens in new tab)).
- Criticism: Aidan Gomez criticized the campaign as moral arrogance (@aidangomez (opens in new tab)). Lucas Beyer noted a transcript wording change from “create” to “train” (@giffmana (opens in new tab)).
- Lobbying: Olah’s team reportedly lobbied the Pope’s advisers to take model consciousness seriously. He ultimately attended (@kimmonismus (opens in new tab)).
- Anti-safety influence campaign: A report describes a group planning to spend at least $100M, run by a former White House deputy chief of staff, that frames AI warnings as a coordinated campaign (@NeelNanda5 (opens in new tab)).
- New organizations: Nathan Lambert and Tom Zick launched Trillium Labs, a non-profit for open post-training recipes and infrastructure (@natolambert (opens in new tab)).
- Funding: Initial support comes from Halcyon Futures and Schmidt Sciences.
- Underdog: The private on-device AI startup announced backing from a16z, Khosla and others (@0xSigil (opens in new tab)).
- Funding: Initial support comes from Halcyon Futures and Schmidt Sciences.
- Governance and markets: Yoshua Bengio joined Canada’s new National Council on AI (@Yoshua_Bengio (opens in new tab)).
- Meta: Meta has parted ways with Virtue AI (@AndrewCurran_ (opens in new tab)).
- Nvidia: Bloomberg reports a record high near $5.7T market value after a $150B buyback increase (@kimmonismus (opens in new tab)).
- Meta: Meta has parted ways with Virtue AI (@AndrewCurran_ (opens in new tab)).
Top tweets (by engagement)
- NYT report on Olah and the Pope’s encyclical (opens in new tab) — 28.7K
- Karpathy’s “land or water” eval (opens in new tab) — 16.4K
- Claude Code “You should know” plugin (opens in new tab) — 7.3K
- Altman on dots (opens in new tab) — 6.8K
- Muse Gadgets announcement (opens in new tab) — 4.9K
- DeepSeek Harness desktop builds (opens in new tab) — 4.3K
- Altman on the Cerebras partnership (opens in new tab) — 4.1K
- Trillium Labs launch (opens in new tab) — 2.6K
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. Qwen Local Inference: 27B Benchmarks, Fine-Tunes, and MTP
- I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. (opens in new tab) (Activity: 1192): OP built backburner, a
llama.cppfork / distributed inference setup (opens in new tab) that offloads part of Qwen 3.8 27B IQ4_XS from a24 GBM4 Pro MacBook to an iPhone 17 Pro Max over10 Gb/sUSB-C: the Mac runs layers1–40, streams activations, and the phone runs layers41–64using Metal 4 tensor ops. Reported end-to-end prefill gains vs Mac-only were+35%at8k,+44%at16k,+29%at32k, and+30%at48k; a cold27ksession improved from245 sstockllama.cpp/228 sfork Mac-only to168 swith the phone. Above64kcontext, the phone instead hosts old KV pages—up to roughly5.7 GB, enabling196k–229k8-bit context allocation—and computes old-key attention, with a140kcontext test improving generation latency from279 ms/tokento176 ms/tokenwhen adding Neural Engine-compiled16kkey pages.- A technically relevant follow-up asked whether the same iPhone-as-secondary-GPU approach could extend to iPads, especially higher-end iPad Pro configurations with more capable Apple Silicon and potentially more RAM. The implication is that iPads might provide better offload performance or hold a larger portion of the context window than an iPhone, making them a stronger companion device for local LLM inference.
- Qwen3.8-27B-Humanlike-Chat 2.0: texts like a human, now with tool calls and better instruction following (opens in new tab) (Activity: 805): LessThanThreeAI released Qwen3.8-27B-Humanlike-Chat 2.0, a merged LoRA over huihui-ai’s abliterated Qwen3.8-27B, available as GGUF/BF16/LoRA on Hugging Face (opens in new tab) with a demo Space (opens in new tab). v2 replaces plain SFT with on-policy distillation: the student generates replies while two teachers score tokens—v1 + hidden “text like a person” instruction for chat/character behavior, and the base model for instruction-following, tools, and code—improving tool-use and controllability while preserving informal texting style. Reported evals vs the abliterated base: IFBench
37.3 → 43.7, When2Call48 → 58, BFCL irrelevance60 → 78, ties/slight gains on IFEval/GSM8K/BFCL simple (83.5 / 89.1 / 98), but regressions on MMLU-Pro (78.5 → 72.5) and LiveCodeBench (56 → 51); a custom “ishuman” judge benchmark rated it as human-written23.5%vs0.3%for the abliterated base and15.1%for official Qwen3.8-27B. Technical discussion in the top comments was sparse; the only relevant critique was that the model’s “humanlike” register may read more like teenage texting than broadly human conversation.- A commenter raised a model-transfer question: whether the same humanlike chat fine-tuning/alignment method used for Qwen3.8-27B-Humanlike-Chat 2.0 would produce similar results on Gemma 4 31B. This is the only technically substantive thread, touching on cross-architecture generalization of the training recipe and whether behavior-style tuning would carry over to a larger Gemma-family model.
- The gap is smaller than they told you: local 27B nearly matches frontier on real code tests (opens in new tab) (Activity: 730): OP reports a single-task DeepSWE/local-code benchmark run using Qwen3.8-27B GGUF via llama.cpp b11115 + llama-swap v257 on 1× RTX 4090 24GB, specifically
Qwen3.8-27B-UD-IQ4_XS.gguf(14.25GB) from unsloth/Qwen3.8-27B-GGUF (opens in new tab), atctx-size 196608,IQ4_XS,q8_0K/V cache, speculative MTP draft, and DeepSeek-style reasoning budget4096. Measured results:115 tok/sdecode,22,934 MiBpeak VRAM,12/12on a code-review task, and on one DeepSWE task40/43hidden tests plus109/109existing tests, i.e. partial0.980but binary pass0; OP later corrected the comparison: the cited96.6%was mean partial across all published trials, while the frontier subset for that task was99.8%partial and85.3%pass, from DeepSWE v1.1 raw data (opens in new tab). The linked writeups cover the 24GB fit/context setup (context ceiling (opens in new tab)) and the task-level DeepSWE result (local confidence (opens in new tab)); OP emphasizes this is task-specific, not a claim that a 27B local model matches frontier models broadly across the113-task benchmark. Commenters were skeptical of the broader framing: one user with both “Flash and 27B” said “the gap is real,” and another argued the conclusion is wrong because even frontier coding models are uneven and~24–72Bmodels may handle discrete subtasks but often lose value once humans must decompose larger engineering work into model-sized tasks.-
Several commenters argued the claimed near-parity is likely an artifact of a saturated benchmark: a local
27Bmodel, even atQ8, can perform well on small/discrete coding tasks but still fails on harder real-world tasks requiring frontier models such as Claude Opus.-
A recurring technical objection was that coding evaluations often underweight project-level decomposition:
24B–72Blocal models may solve isolated tickets, but for larger work items the human effort needed to break problems into model-sized subtasks can exceed the productivity gains. -
Users with hands-on experience running both Gemini Flash and local
27Bmodels reported that the performance gap remains substantial, especially for nontrivial coding workloads where frontier models provide better reliability and task completion.
-
A recurring technical objection was that coding evaluations often underweight project-level decomposition:
-
Several commenters argued the claimed near-parity is likely an artifact of a saturated benchmark: a local
- Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp (opens in new tab) (Activity: 405):
llama.cppPR #29761 (opens in new tab) adds MTP speculative decoding support for Qwen3.8-Flash Next via--spec-type draft-mtp, merged into theaman/qwen4-optbranch after ~17hof development. Reported DGX Spark benchmarks for Qwen3.8-Flash-Nextiq4_xswith-np 1 -lzm on --spec-draft-n-max 3show decode throughput improving from28.36to43.88 tok/s(1.55×), latency speedup of 1.54×, and mean speculative acceptance of0.640across24tasks; GGUF quants are available on Hugging Face (opens in new tab). Commenters noted the model is still impractically large for many local setups: theIQ4_NLGGUF is split into a tiny10.9 MBshard plus a102 GBshard, undercutting the idea of casually switching from Qwen 3.8 27B. One commenter also noted that Gufo supports MTP.-
One commenter reported that enabling MTP made inference slower in their testing, arguing it may be more useful for dense models than very large overall architectures where the MTP head has a low acceptance/hit rate. Their hypothesis is that the MTP head cannot effectively predict/compress enough of the larger model’s behavior, reducing speculative decoding benefit.
- A user noted that Gufo already supports MTP, implying llama.cpp is catching up with existing MTP-capable tooling/backends for Qwen-style experimental models.
- Another commenter said they had been using an EXL3 version through tabbyapi because llama.cpp GGUF inference was “way, way slower” for their workload. They planned to retest after this PR, but their prior experience suggests EXL3/tabbyapi may still be a performance baseline to compare against for Qwen4Exp/MTP support.
-
One commenter reported that enabling MTP made inference slower in their testing, arguing it may be more useful for dense models than very large overall architectures where the MTP head has a low acceptance/hit rate. Their hypothesis is that the MTP head cannot effectively predict/compress enough of the larger model’s behavior, reducing speculative decoding benefit.
2. Local Agent Tooling: Decision Models and MCP
- Pi 1.0 released - MCP support now included by default (opens in new tab) (Activity: 679): Earendil released
Pi 1.0, a stable version of its minimal agent harness, with Codemode now including native MCP support by default plus non-LLM/image model support, virtual-model extensions, deferred tool loading, Anthropic cache warming, mid-conversation system messages, and TUI updates. The release also introduces experimental MIT-licensed Pi Durable for longer-running agentic applications beyond terminal/coding-agent workflows, while retaining Pi’s minimal/extensible architecture. Top comments focused on naming ambiguity— “pi” collides with many AI/dev tools—and requested clarification of what Codemode is. One commenter linked Earendil’s rationale for MCP support: “You said no MCP” (opens in new tab).- A commenter linked the maintainer’s rationale for reversing course on MCP support in Pi 1.0, pointing to the post “You said no MCP” (opens in new tab). The thread notes that MCP is now included natively/by default, after earlier resistance from the creator based on project ethos, with users framing it as a “vital addition” for tool/server integration workflows.
- Clef: Open Weights decision model by Cloudflare (opens in new tab) (Activity: 619): Cloudflare announced Clef, an open-weights “decision model” intended for local/self-hosted use. A top commenter notes that Clef was post-trained from Qwen3.8-27B and that clef-flash was also released, post-trained from Qwen3.5-9B. The main substantive reaction was positive: commenters see Clef as filling a gap in the local-model ecosystem and are eager to benchmark it themselves.
-
Commenters noted that Cloudflare Clef is post-trained from Qwen3.8-27B, with a smaller clef-flash variant post-trained from Qwen3.5-9B, framing it as a potentially important open-weights “decision model” for local inference use cases.
- A technical concern raised was how Clef’s quality holds up after quantization, especially below Q8, since local deployment will likely depend on lower-bit quantized variants and decision-model behavior may degrade nonlinearly under aggressive compression.
- The benchmark discussion focused on comparisons against models such as Laya, Kev 9B, and DiffusionGemma Jev, but one commenter criticized the eval set as too weak and argued Clef should be compared against the leading models on jevbench rather than weaker open Jev baselines.
-
Commenters noted that Cloudflare Clef is post-trained from Qwen3.8-27B, with a smaller clef-flash variant post-trained from Qwen3.5-9B, framing it as a potentially important open-weights “decision model” for local inference use cases.
- New in llama.cpp: Decision Models (opens in new tab) (Activity: 574): The post announces Decision Models support in
llama.cpp: local “Jev/Jeff-like” models intended to act more like controllers/classifiers—selecting among actions, continuations, or behavioral choices—rather than purely free-form generators. No benchmark numbers or low-level implementation details were discussed in the provided comments; the main concrete use case raised was steering local roleplay models to avoid characters “going off the rails mid scene.” Commenters were skeptical about Jev as a defensible product/category, arguing the idea had “no moat” and was rapidly cloned into many Jev-like models. Others said they still do not know what these models are practically useful for, aside from possible agent/roleplay control.-
A commenter frames decision models as essentially a constrained classification loop: provide a JSON schema containing allowed classes plus structured/unstructured input, then have the model select exactly one category per item. The technical question is whether llama.cpp’s new support adds meaningful inference-time behavior beyond ordinary prompt-constrained JSON classification, or primarily standardizes the workflow for local models.
- One practical use case raised is applying decision models to local roleplay agents to reduce derailment during long scenes—i.e., using an auxiliary model or decision step to enforce state/intent constraints before generation. The thread does not report benchmarks or implementation results, but highlights a potential control-layer pattern for character consistency and scene-state management.
-
A commenter frames decision models as essentially a constrained classification loop: provide a JSON schema containing allowed classes plus structured/unstructured input, then have the model select exactly one category per item. The technical question is whether llama.cpp’s new support adds meaningful inference-time behavior beyond ordinary prompt-constrained JSON classification, or primarily standardizes the workflow for local models.
Less Technical AI Subreddit Recap
/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo
1. Gemini 4 Argon Access Backlash
- Just canceled my Google One AI plan. (opens in new tab) (Activity: 1838): The OP claims Google’s paid Google One AI Pro tier no longer provides access to frontier Gemini models: after an alleged Gemini 4 Argon announcement, access is described as limited to enterprise “Fairwind” partners, paid API users, and a forthcoming Google AI Ultra tier, while Pro users remain on Gemini 3.8 Flash. They argue this is a regression from the Gemini 2.5 Pro era—where higher-end reasoning models and generous limits were available more broadly—and contrast it with Anthropic/OpenAI subscriptions allegedly offering frontier models to standard paid users; no concrete benchmark numbers are provided beyond claims of “impressive benchmark charts” and a
1Mtoken output ceiling. Top comments mostly dismiss the complaint: one user says they subscribe primarily for Google storage and treat AI as a bonus, while others question the post’s authenticity, alleging it was Gemini-written or bot/shill activity from a new account.-
One commenter argued that the
$20/monthAI subscription tiers from Google/Anthropic/OpenAI function more like constrained trials than production-grade access, implying practical limits on sustained workloads despite “Pro” branding. They also suggested Google may be subsidizing or losing money on Google One AI Pro subscriptions given the underlying inference costs.- A rollout clarification noted that Gemini Ultra appears to be receiving access first, but Google has not explicitly ruled out Pro-tier access to features like Astra. The commenter framed this as a typical staged software rollout rather than definitive permanent tier exclusion.
-
One commenter argued that the
- Why publicly announce a model that the public can’t use yet?? (opens in new tab) (Activity: 1624): The image is a screenshot of a purported Google/Gemini announcement for “Gemini 4 Argon”, claiming frontier performance in software engineering, knowledge work, and cybersecurity defense, plus an extremely large
1M token output limit: image (opens in new tab). The post’s technical significance is mainly about model-release communication, not evaluation: the title questions why Google would publicly announce a model before it is accessible to users, and the image itself provides no benchmarks, API details, pricing, or availability timeline. Commenters compare this to prior “announced but unavailable” model rollouts, including Anthropic’s “mythos” and Google’s alleged “3.5 pro” handling. The dominant view is skeptical: users may be frustrated, and one commenter speculates the announcement is aimed “purely for investors.”
2. Claude Opus 5.5 Regression Reports
- Opus 5.5 nerfing - how to measure, how to spot, how to sue (opens in new tab) (Activity: 2722): Poster alleges Anthropic Opus 5.5 showed a sharp post-launch regression after
5–6days on complex C++/3D/physics/Blender MCP workloads, citing abnormal phrasing and lower code/output quality, and recommends preserving exact launch-day prompts/outputs plus latency measurements to detect potential changes such as quantization, routing, or serving optimizations under load. They frame this as a potential EU consumer-law issue under the Digital Content Directive 2019/770 (opens in new tab), specifically conformity expectations in Arts.7–8and modification/withdrawal notice obligations in Art.19, arguing launch benchmarks and “most capable model” marketing may set enforceable expectations. Comments broadly agree that closed-model providers can silently degrade or reroute models and that independent auditing is needed, but no commenter provides reproducible benchmarks or direct evidence. One commenter reports similar perceived quality drops in Higgsfield outputs, describing wasted credits after initially strong generations.-
Commenters raised the core measurement problem with alleged closed-model degradation: because Anthropic’s hosted model weights, prompts, routing, and serving configs are opaque, users argue it is difficult to prove a regression or “nerf” without independent auditing, fixed benchmark prompts, repeated sampling, and historical baselines. One user specifically asked for an “effective test or reliable nerf tracker site,” highlighting demand for third-party longitudinal evals rather than anecdotal comparisons.
- Several users reported anecdotal regressions in Opus 5.5 behavior across applied workflows: one claimed it now needed help from Gemini 3.8 Flash to catch coding bugs, while another said Higgsfield design/render outputs declined after initially strong results, wasting credits. These reports are not controlled benchmarks, but they point to the kinds of tasks users want tracked: bug-finding accuracy, design/render prompt fidelity, and day-over-day output consistency.
-
Commenters raised the core measurement problem with alleged closed-model degradation: because Anthropic’s hosted model weights, prompts, routing, and serving configs are opaque, users argue it is difficult to prove a regression or “nerf” without independent auditing, fixed benchmark prompts, repeated sampling, and historical baselines. One user specifically asked for an “effective test or reliable nerf tracker site,” highlighting demand for third-party longitudinal evals rather than anecdotal comparisons.
- Mmmkay. I didn’t believe others at first, but something is suddenly off with Opus 5.5 (opens in new tab) (Activity: 2045): A Claude Code Enterprise PAYG user reports a sharp perceived regression in Claude Opus 5.5 Med behavior after a monthly limit reset: from architecture-first, DRY/SOLID, token-efficient implementation to verbose preambles, duplicated code, “slopcode,” and token burn resembling prior Opus 5 behavior. They claim usage jumped from roughly
70%to90%in about an hour, versus no spend-limit increase requests during the previous week of heavy~12h/dayO5.5 use, and offer daily cost/token data for comparison. A commenter cites external sentiment tracking showing Opus 5.5 Reddit sentiment dropping from71–73/100on Sep 25–28 to58yesterday and55today on modelsentiment.com (opens in new tab), while noting it measures opinion rather than backend model changes. Top comments speculate Anthropic may have reduced compute, silently changed routing, or altered token accounting after launch hype, but no direct evidence is provided. The main debate is trust/reliability: users want stable model behavior and transparent deployment/versioning rather than perceived post-release regressions.-
A commenter tracking Reddit sentiment reports a sharp drop for Claude Opus 5.5, with scores allegedly stable at
71–73/100from Sep 25–28 before falling to58yesterday and55today on modelsentiment.com (opens in new tab). They note this measures user opinion rather than model behavior, so it cannot confirm a backend change, but it may indicate a sudden perceived quality regression.- Multiple users describe a suspected capability regression in Opus 5.5, especially around instruction-following and multi-part prompt adherence: one says the model now “mentions 3 things and only acknowledges 2,” and even recognizes the omission when challenged. Another user says they reverted to “xhigh effort” mode for all tasks, implying lower default reliability or reduced reasoning/compliance under normal settings.
- One technical hypothesis raised is that Anthropic may have temporarily allocated more compute during launch/benchmarking and later reduced inference resources or altered token accounting, leading to perceived quality degradation. This is speculative and unverified, but the complaint centers on reproducibility and reliability: users want model behavior to remain stable after release rather than changing silently under the same product name.
-
A commenter tracking Reddit sentiment reports a sharp drop for Claude Opus 5.5, with scores allegedly stable at
3. AI Video Models and Motion Control
- Orbiting Lora + first and last frame in MiniMax gives fantastic results (opens in new tab) (Activity: 2263): A user shared a MiniMax-H3 LoRA for generating locked-subject
360°orbit shots from first/last-frame conditioning:pablodawson/MiniMax-H3-360-Orbit-LoRA. The prompt explicitly constrains the scene to a frozen instant—no object/pose deformation, no drifting, no continued action—so that camera parallax is the only motion source, targeting cleaner pseudo-volumetric outputs suitable for downstream reconstruction workflows. A linked Reddit demo video was mentioned, but the video URL could not be inspected due to Reddit returning403 Forbidden. Commenters framed the LoRA as especially useful for creating 3D assets: one suggested feeding the generated orbit clip into Opus to extract snapshots for 3D model generation, claiming it improves style preservation. Another commenter extrapolated that this kind of orbit-consistent video generation brings consumer volumetric/VR viewing of existing films closer.-
One commenter describes a workflow where an orbiting/generated 3D video snippet is fed into Opus and instructed to extract snapshots for 3D model creation. They report that this improves style capture substantially versus prompting the model without the video reference, suggesting the orbit video acts as a strong multi-view conditioning source.
- A technical artifact noted in the output is inconsistent motion segmentation: humans remain effectively frozen while secondary elements such as the car, hair, and background explosion continue moving. This points to MiniMax preserving the subject pose from the first/last-frame constraints while still synthesizing environmental dynamics, which can create partial-animation mismatches.
-
One commenter describes a workflow where an orbiting/generated 3D video snippet is fed into Opus and instructed to extract snapshots for 3D model creation. They report that this improves style capture substantially versus prompting the model without the video reference, suggesting the orbit video acts as a strong multi-view conditioning source.
- Griffin, the first Human Interaction Model to pass video Turing Test it’s already #1 on NVIDIA’s benchmark for full-duplex AI video - 44% of people thought it was a real person while other systems are at ~3% (opens in new tab) (Activity: 2018): A Reddit post claims Griffin, described as a “Human Interaction Model,” is the first system to pass a video Turing Test and ranks
#1on NVIDIA’s benchmark for full-duplex AI video, with44%of participants judging it as a real person versus roughly~3%for other systems. The linked Reddit video could not be independently accessed due to a 403 Forbidden response, so the benchmark details, methodology, and model architecture are not verifiable from the provided source. Comments were mostly non-technical: one user joked about the human/AI reveal being reversed, while another argued the technology is unnecessary and likely to be used in predatory applications.-
A commenter emphasized that a
44%human-identification rate is technically significant because humans are usually highly sensitive to subtle facial, timing, and behavioral anomalies—the basis of the uncanny valley problem in CGI/animatronics. They argued this suggests Griffin is substantially beyond prior “fake human” systems, especially compared with the post’s claim that other systems score around~3%on the same video Turing-style benchmark.- One technically relevant real-world abuse case raised was AI-generated job applicants: synthetic candidates allegedly apply, conduct video interviews, get hired, and then either gain internal platform access or steal shipped work equipment. The commenter noted that large companies with weak scrutiny or limited background checks may fail to detect these AI-mediated interviews, implying full-duplex video agents could materially worsen identity-verification and hiring-security risks.
-
A commenter emphasized that a
- OpenAI priced GPT-6.1 Sol at $2/$10 per million input/output tokens, versus $10/$50 for Astra; reported results put Sol 6.4 points above GPT-6 Sol on DeepSWE v1.1 and 2.2 points above Opus 5.5 on AutomationBench. In Agent Arena, Sol [Max] ranked #5 at $0.56 per task—39% cheaper than GPT-6 Sol while scoring 1.52 points higher, and 81% cheaper than Astra while within 1.04 points. Sonnet 5.5 debuted at #3 overall and #1 in Chat, but its $2.74 per-task cost exceeded #2 Opus 5.5’s $1.58; Anthropic models held the top three spots.
- Agent results are sensitive to training and setup: the same weights scored 62% in one harness and 33% in another; a multi-harness training effort raised LFM2.5-2.6B from 42% to 54% across four harnesses and cut tool calls by 31%, with its trainer, data, and seven models released openly. A dedicated controller reportedly raised GPT-5.5’s ProgramBench score from 63.7% to 71.5% using the same workers and budget, while AgentWorld found that fewer than a third of multi-agent actions helped and coordination tasks reached only 12% success.
- Epoch estimates that infrastructure could support hundreds of millions to billions of agents; at 20% utilization, its scenario implies $2.6–5.3T in annual spending versus roughly $1T in lab revenue by end-2027. These are projections, not observed demand. Optics startup Volantis is targeting up to 10K tokens/second per user on models over 10T parameters.
- Nathan Lambert and Tom Zick launched Trillium Labs, a nonprofit focused on open post-training recipes and infrastructure, with initial support from Halcyon Futures and Schmidt Sciences. Private on-device AI startup Underdog announced backing from a16z, Khosla, and others.