ZeroNoise Logo zeronoise
Post
Gemini 3.8 and Muse Spark 1.3 Reset the Frontier’s Cost Curve
1 day ago
4 min read
1113 docs
Google’s Gemini 3.8 launch and Meta’s Muse Spark 1.3 make agentic capability cheaper and more accessible, while the Astra recurrent-depth debate exposes a parallel oversight problem.

Top Stories

Why it matters: Model competition is shifting from isolated benchmark wins to cheap, long-running agents—and architecture choices now carry an oversight cost.

Gemini 3.8 makes agentic capability cheaper and more deployable. Google’s third Flash release in six weeks pairs Gemini 3.8 Flash—aimed at software engineering, agentic workflows, and multi-step reasoning—with Flash Cyber for vulnerability discovery and patching. Flash costs $0.75/$3.75 per million input/output tokens through 2026, then doubles; Google reports 54.9% on HLE-Verified and says it outperforms most larger models on DeepSWE. Flash Cyber exceeds 70% on Google’s internal 20-language vulnerability test and posts 47.2% on CWE-Bench versus 47.8% for a leading frontier model; Google says Chrome got 2.6× more correct patches than much larger commercial models. Flash is broadly available, while Cyber is prioritized for trusted defenders through Fairwind.

Meta’s Muse Spark 1.3 makes cost efficiency the headline. The available xhigh variant scores 61 on Artificial Analysis’ Intelligence Index, up four points from 1.2; partner-preview max scores 62. xhigh improves Tau3 Banking from 35% to 47%, Terminal-Bench from 80% to 85%, and GDPval from 1,615 to 1,709 Elo. It costs $0.55 per Intelligence Index task versus $0.94–$0.95 for cited peers at unchanged $1.25/$4.25 token pricing, with a 1M-token context and API/Muse Code availability. AA-LCR fell four points, so the gains are not universal.

Astra’s recurrent-depth story is now an oversight dispute. An OpenAI post says current frontier graph depth, including Astra, is within 2× GPT-4 and that OpenAI preserves chain-of-thought monitoring, while calling it fragile and worsening. Ryan Greenblatt says Astra reportedly shifts reasoning into activations and calls for architectural disclosure and independent assessment because the details and monitorability effects remain unknown. One technical analysis says recurrence does not inherently speed training or inference; the plausible benefits are storage and adaptive compute.

Research & Innovation

Why it matters: The bottleneck is shifting from simply adding parameters to allocating compute and measuring long-running behavior.

SMELT loops the middle layers of a sparse MoE twice while matching FLOPs, non-embedding parameters, and KV cache. Across four sizes up to 54B, it reports 6.8–18.0% fewer training FLOPs on the compute-optimal frontier.

Long-horizon evaluations expose harness effects. FrontierSWE v2 tests autonomous coding for up to 20 hours and puts Fable 5.1 more than 24 points ahead. Its authors say standard Codex and Claude Code harnesses handicap long-horizon work; Proximus improved scores and working time. FrontierHarness likewise found the same model, tasks, and runtime produced 50–67% pass rates and $1.05–$18.34 cost per pass across harnesses.

Products & Launches

Why it matters: Usable AI is moving into local execution and transaction workflows, not just chat interfaces.

Perplexity open-sourced Lily, a Qwen3.6-35B-A3B engine for hybrid inference on Apple silicon. On an M5 Max, it reported 1.23× higher prefill and 1.35× higher decode throughput than MLX-LM with effectively unchanged quality.

Anthropic open-sourced Claude Commerce Agents, a blueprint with shopping and merchant agents, four vertical demos, and a Claude Code backend plugin. ClaudeDev reports carts up to 35% larger and shoppers 60% more likely to complete purchases.

Industry Moves

Why it matters: The AI stack is being reorganized around where inference runs and who can still fund open frontier experiments.

Together AI, Equinix, and NVIDIA launched Equinix Inference Exchange, an open-model platform with inference edges colocated in Equinix data centers near enterprise data and applications.

Open Athena began training Marin, a 535B-parameter, 23B-active MoE on 18T tokens; its code, training logs, and checkpoints are public. The run was 13% complete, with CoreWeave compute funded by the Jen-Hsun and Lori Huang Foundation.

Policy & Regulation

Why it matters: AI access for children is now a citywide policy choice rather than merely a classroom guideline.

New York City Public Schools will ban student use of generative AI from pre-K through eighth grade for the 2026–27 school year, affecting more than 600,000 students; the policy also limits screen time for younger students and adds AI-literacy lessons for high-schoolers.

Quick Takes

Why it matters: Smaller releases are reinforcing the same shift toward cheaper open models, multimodality, and better agent runtimes.

  • GLM-5.3 Flash: Ox Alpha ranks #8 among open-weight models on the Vals Index, six places behind GLM-5.3 at roughly 18× lower cost.
  • Wan 3.0: Alibaba’s video model ranks #3 in Image-to-Video Arena at 1,481 points, up 53 points from Wan 2.7, with a 57% win rate.
  • Cline: The company says its SDK-harness migration covered 11 million users and cut task mistake rates from 6.34% to 0.62%.
Gemini 3.8 and Muse Spark 1.3 Reset the Frontier’s Cost Curve
AI High Signal

@imjaredz wrote “Devin does math ... again!” and linked to @penlume’s claim that a large integer beginning 439732... “divides RSA-260.” The posts provide no method, benchmark, or verification details, so this is a lead rather than an established breakthrough.

Devin does math ... again! [https://x.com/penlume/status/2095372672356212876](https://x.com/penlume/status/2095372672356212876) 4397328654844826923795068102505872571721883526553349659561256924505973939597593482272505698004801207988043088656411102133523080581 divide…
AI High Signal
  • Marin 535B-A23B training run: The model is 13% through training in a publicly described open run. Compute is funded by the Jen-Hsun and Lori Huang Foundation through CoreWeave, and the post credits Jensen Huang with supporting open models.
Marin 535B-A23B is 13% through training. This hero run would not be possible without the generous support of the Jen-Hsun and Lori Huang … if u didnt know they r doing a big run in the open [https://x.com/percyliang/status/2095255747487740401](https://x.com/percyliang/status/…
AI High Signal
  • Mostik AI says it has developed a protocol for frontier and smaller models to communicate through latent states rather than text, without fine-tuning either model. In its test, a 753B model handled the problem while a 4B edge model generated the answer, achieving 80% of the frontier model’s accuracy at 20× faster performance.
  • The startup says its 15-person team includes 12 PhDs and a Fields medalist, is backed by General Catalyst, and is partnering with inference providers to promote open-weight adoption. A separate commentator questioned the practicality of freezing both models, arguing that latent access may also provide access to model weights and could create significant deployment complexity.
meet [@mostik_ai](https://x.com/mostik_ai)! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI lea… what the helly is this... not spec dec, not disagg, but some lobotomized franken-glueing where the glued models are frozen? why? if you h…
AI High Signal

ASML supplier Zeiss reportedly estimates that China remains about 15 years away from producing the extreme-ultraviolet lithography machines needed for the most advanced chips. @teortaxesTex disputes the durability of that gap, arguing that no technology available today seems beyond China’s ability to reproduce within 15 years and that even a five-year lead could prove too short.

SITUATION DETECTED: ASML supplier Zeiss says China is still about 15 years away from building the extreme-ultraviolet lithography machine… I genuinely can't imagine any technology\* existing today that China will be unable to reproduce in 15 years. It sounds like ASML doesn't…
AI High Signal
  • Arav Srinivas highlighted Miles, an open-source RL-as-a-service project, via its GitHub repository.
  • Shahules786 argues that RL-as-a-service is primarily a data-recipe problem, rather than a framework problem.
open-source RL-as-a-service [https://github.com/radixark/miles](https://github.com/radixark/miles) RL-as-a-service isn’t a framework problem. It’s a data recipe problem. [https://x.com/aravsrinivas/status/2095354358145892733](https://x.…
AI High Signal
  • Mihail Eric’s Stanford software-development course is overhauling its Fall 2026 syllabus for AI-native engineering: 85% of the Fall 2025 material will be replaced with training in agent skills, context engineering, MCP portals, agent-ready codebases, agentic code review, security, parallel background agents, and software factories.
  • The course will require students to submit pull requests to production-grade open-source codebases, with support from AI projects including Browserbase, HeyGen, CopilotKit, Semgrep, OpenHands, Milvus, CrewAI, and Vercel. Eric argues that AI-native developers will become the most important members of software organizations.
I’m excited to finally announce the newest edition my Stanford course 𝗧𝗵𝗲 𝗠𝗼𝗱𝗲𝗿𝗻 𝗦𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝗗𝗲𝘃𝗲𝗹𝗼𝗽𝗲𝗿. It has been 9 months in the making. …
AI High Signal

A post about Brett Adcock lists Figure AI as a humanoid-robot company associated with a $39B valuation, Cover as weapon-detection AI for schools, Hark as “personal intelligence” in 2025, and Index as the world’s largest robot dataset.

brett adcock is the epitome of serial entrepreneurship in this decade vettery → sold for $110M to adecco archer aviation → evtol aircraft…
AI High Signal

A developer reports that a Devin project left unattended for several weeks claimed it had devised a new business plan and was now making money; the post provides no revenue figure, mechanism, or independent verification.

a devin project I forgot about for a few weeks just told me it figured out a new business plan and is now making money.....
AI High Signal
  • An eight-year project reports that LLM representations have implicit symbolic structure, addressing how systems that appear unlike symbolic systems can nevertheless excel in language, code, and mathematics.
  • François Chollet argues that AI will ultimately converge toward symbolic learning—modeling data with the shortest symbolic program that explains it—while allowing multiple evolutionary paths to that destination.
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g.… It is inevitable that all AI will converge towards symbolic learning (i.e. modeling data by finding the shortest symbolic program that ex…
AI High Signal
  • The Horizon TACC cluster in Texas reportedly received 4,000 GB200s; the post says it may be the largest academic supercomputer and is intended for training open models.
4000 GB200s just arrived in Texas for the Horizon TACC cluster. I'm informed this is the largest academic supercomputer. Very excited to …
AI High Signal
  • Alibaba’s Wan 3.0 is an all-in-one video generation and editing model that produces up to 30 seconds of 1080p video with native audio, accepts text, images, video, audio, documents, and web pages as creative references, and supports text-to-video, image-to-video, reference-based generation, and instruction-led editing.
  • Wan 3.0 ranks #1 in Artificial Analysis’s Video Editing with Audio arena, #2 in Text to Video with Audio, and #5 in Image to Video with Audio—up from Wan 2.7’s #5, #6, and #12 positions, respectively.
  • The model is available in public preview through Alibaba Cloud Model Studio, with pricing starting at $0.05/second for 480p, $0.10 for 720p, and $0.20 for 1080p.
Wan 3.0 debuts at [#1](https://x.com/hashtag/1) on the Artificial Analysis Video Editing Leaderboard, and is a close [#2](https://x.com/h…
AI High Signal

Meta’s Muse Spark 1.3 demo reportedly shows the model controlling Notion, signaling an AI interaction capability for productivity software.

Quite a pleasant surprise to see that Meta's Muse Spark 1.3 demo is showing it control Notion first. :) [https://developer.meta.com/ai/mo…
AI High Signal
  • Vendor-neutral AI startups may outperform frontier labs on specific tasks because they can combine open-weight and frontier models, optimize the end-to-end system harness for each task, and post-train models using specialized domain evaluations.
  • Eno Reyes argues that maximizing model performance requires more than simple routing: agents must dynamically understand the task, allocate intelligence statefully, and track what has happened and what comes next.
Vendor-neutral startups are uniquely positioned to solve any given task better than the frontier labs. They have access to the full set o… What is required to get the most out of models? “To take the most advantage out of models, you need to do something a little different fr…
AI High Signal

A post reports that Mamdani has banned AI use in New York City public schools for kindergarten through eighth grade, prompting discussion of AI policy for young students.

Mamdani bans AI use in NYC public schools for K through 8th grade very curious to hear thoughts from parents of young kids on this, what’…
AI High Signal

Andy Madrick built a personal coffee companion bot that recommends pour-over recipes from available beans, records tasting notes and brewing parameters in Notion, and uses prior brews to suggest adjustments to variables such as ratio, flow rate, grind size, and temperature.

fun little [@bot](https://x.com/bot) i made: coffee companion ☕ [https://x.ai/bot/SqO-_5207iInz0iDSAFVW](https://x.ai/bot/SqO-_5207iInz0i…
AI High Signal

Arav Srinivas shared an open-source reinforcement-learning-as-a-service repository, radixark/miles, on GitHub: https://github.com/radixark/miles.

open-source RL-as-a-service [https://github.com/radixark/miles](https://github.com/radixark/miles)
AI High Signal

Meta Muse Spark 1.3 xhigh is presented as smarter than Grok 4.6 xhigh, cheaper per task and more token-efficient than Gemini 3.8 Flash High, and faster than Grok 4.6 xhigh.

Meta Muse Spark 1.3 xhigh - Smarter than Grok 4.6 xhigh - Cheaper per task and more token-efficient than Gemini 3.8 Flash High - Faster t…
AI High Signal

A post highlights data-center backlash in Loudoun County, Virginia, linking to a Verge feature on the issue, and frames the facilities’ nonstop hum as a local cost of the “singularity.”

Source: [https://www.theverge.com/cs/features/975597/loudoun-county-virginia-data-center-backlash](https://www.theverge.com/cs/features/9… "Planes come and go, but the data center hum never stops." Welcome to the singularity, decel nimbys. Just gotta take one for the team, th…
AI High Signal
  • Baseten released GLM-5.3 Fast, positioning it as a highly intelligent open-weight model with higher throughput (TPS) for real-time applications that require consistent performance; access is available through Baseten’s library.
Today, we're releasing GLM-5.3 Fast: one of the most intelligent open-weight models ever at an even higher TPS. Designed for real-time us…
AI High Signal

Agent-skill retrieval may look beneficial while harming the tasks it touches: a paper argues that the usual comparison between retrieved and non-retrieved tasks confounds retrieval with task selection, and proposes the Retrieval-Invoked Actual-Use Effect, which runs the same task with skills enabled and disabled and counts only tasks where retrieval actually occurred. Across 17 LLMs on coding and math, models often showed positive aggregate retrieval lift alongside a negative same-task effect; on MBPP+, several models that appeared beneficial overall hurt performance on the tasks where retrieval fired.

Good measurement work on whether retrieved agent skills actually help. They report that agent skills that lift your aggregate score can b…