ZeroNoise Logo zeronoise
Post
Jev’s Control Layer Meets Muse’s Consumer-Agent Push
4 min read
712 docs
A concise intelligence brief on Jev’s evaluation economics, Muse’s consumer-agent traction, open model launches, and the infrastructure and governance shifts around agentic AI.

Top Stories

Why it matters: The competitive frontier is shifting from raw model output to systems that make bounded decisions and complete everyday tasks.

Jev is turning bounded judgment into agent infrastructure. TypeSafe describes Jev as a non-generative “System One” model that returns typed answers and probabilities for software to use directly. In LangChain’s narrow test—five weather requests, fixed traces, and 100 repetitions—it matched the human oracle on all 500 binary decisions; its observed score variance was 92–913× lower than the tested LLM judges. The result is observational and still needs validation on other agents and production workflows, but the reported $0.00035-per-judgment cost makes frequent online checks economically plausible.

Muse is making agentic workflows accessible outside the AI bubble. One user report says it found and booked a barber under budget, schedule, review, and haircut constraints without leaving the app. Alexandr Wang says nontechnical users can use Muse agents without knowing CoT, MCP, or CLI. A related product description says its computer layer routes work across 15-plus models, including GPT and Claude, while customized Muse Spark models are co-developed with the harness. Early reports therefore point to orchestration and distribution as important product differentiators alongside model capability.

Research & Innovation

Why it matters: New results are improving the control loop around models—reducing token cost, staging verification, and exposing safety failures that single prompts miss.

  • NVIDIA’s SoL-Pi automates search over agent-harness mechanisms across repository-derived and verifier-driven environments. Four retained mechanisms nearly halve token traffic while matching baselines on GPT-5.6 Sol and Opus 5; on 51 EdgeBench tasks, the authors estimate roughly one-third lower API cost.
  • Google Research’s Stellar Colosseum uses staged proof search, readiness gates, parallel candidate generation, targeted falsification, and section-level feedback. With Gemini 3.1 Pro and 3.7 Flash, it reports 71.0% on TCS-Bench and 218/222 Codeforces solves when given execution feedback.
  • Microsoft’s “capability laundering” result shows a local unaligned model splitting a harmful objective into benign-looking subquestions, querying an aligned frontier model in separate sessions, and recombining the answers. Gemma-4-31B recovered 8/14 failed CyBench tasks after GPT-5.5 consultation; a CBRN-chain score rose from 62.3 to 83.1, supporting session- and account-level monitoring.

Products & Launches

Why it matters: Open releases and narrow, high-frequency tools are becoming easier to deploy, not just easier to demo.

  • Qwen-Image-2.1 is an open-weight 7B unified generation/editing model with native RGBA output, up to 10 reference images, and precise local editing controls. vLLM-Omni added day-one support using a 7.1B DiT paired with Qwen3-VL-8B.
  • DocJev packages Jev for document classification and splitting from natural-language rules; its author claims 6× faster execution than GPT-5.6 Luna at equivalent accuracy, with open-source and OCR/VLM backend options.
  • Cline Desktop reported running more than 6% of all Cline tasks one week after launch, with half of that activity from new users. The figure is self-reported, but it is an early adoption signal rather than only a feature announcement.

Industry Moves

Why it matters: The frontier race is becoming a race to fund and serve high-volume agent workloads reliably.

  • Xiaomi’s MiMo reinforcement-learning run reportedly consumed $3.241 million over 4.5 days: Pro accounted for $2.387 million at about $20,600 per hour and a 70.92 DeepSWE score. The run used 23 harnesses and hit GPU OOM, VRAM, grader-network, and infrastructure failures—a live cost-and-reliability test, not a settled performance result.
  • China’s open-model business model is unsettled. SemiAnalysis’ Dylan Patel says multiple Chinese labs told inference providers that upcoming models would be licensed rather than open-sourced; another commentator, citing conversations with Chinese labs, says they plan to keep open-sourcing and monetize through inference revenue share.

Policy & Regulation

Why it matters: Voluntary model rules are starting to specify prohibited capabilities before formal policy catches up.

Microsoft’s new AI code of conduct reportedly imposes absolute constraints against cyberattacks, nuclear weapons, and deepfakes, while barring deceptive or collusive mechanisms that evade human oversight. It is a company standard rather than government regulation, but its specificity gives the industry concrete boundaries to test against.

Quick Takes

Why it matters: Research volume and architecture bets are accelerating faster than validation standards.

  • ICLR 2027: Denny Zhou reports more submissions than all previous ICLR years combined.
  • Diffusion LLMs: Inception AI CEO Stefano Ermon pitches parallel token generation to remove today’s sequential bottleneck; the post supplies no performance benchmark.
  • Byte-level models: Meta reports that token models lead at low compute but byte models overtake them as compute grows, reaching up to 4% higher asymptotic performance and matching a distilled token model with one-sixth the training data.
Jev’s Control Layer Meets Muse’s Consumer-Agent Push
AI High Signal
  • Posts are circulating a claim that OpenAI has solved—or is close to solving—the next Millennium Prize Problem, with one post referring to an alleged announcement and citing Cédric Villani’s reaction.
  • Elliot Glazer’s counter-analysis says the rumor is more plausibly about substantial progress on a special case of the Hodge conjecture, potentially Hodge for abelian varieties, rather than the full conjecture; he says this interpretation is unconfirmed and warns that insider enthusiasm may be repeating an earlier rumor that Anthropic had solved two Millennium Problems, which he considers nearly certainly false.
Cédric Villani after the announcement of OpenAI’s solution to the Millennium Prize problem: "I was devastated. An atmosphere of the 'end … Appreciate you showing genuine curiosity here unlike most people I've interacted with on this issue. Part of why I started throwing large…
AI High Signal
  • Jerry Liu introduced DocJev, an open-source library for document classification and splitting: users provide natural-language category rules, and it predicts document categories or boundaries between sub-documents. The project claims 6× faster execution than gpt-5.6-luna at equivalent accuracy.
  • DocJev supports liteparse for fast, free document parsing and LlamaParse for more complex documents via a VLM-based pipeline, with the latter adding preprocessing latency.
  • Liu says Jev is intended to make routine business operations lightweight and fast while reserving intelligence-heavy tasks for larger agentic systems; DocJev is narrowly focused on classification and splitting.
Introducing DocJev - a lightning-fast OSS library for document classification and splitting with jev ⚡️ Give a document alongside some nat… it's awesome to see the reception here 🔥 one of the jev's promises is to make a lot of business operations extremely lightweight and fast…
AI High Signal
  • A September 2026 estimate puts closed-model AI revenue at $125–165B ARR versus $11–15B for open models, implying open models represent roughly 6–11% of global AI revenue. The author describes the estimate as preliminary and uncertain, especially for hyperscaler revenue.
  • The open-model estimate was challenged as materially outdated: the analysis reportedly used Q1 or Q2 data even though Together AI token volume grew more than 10× from Q1 to Q2, Fireworks grew 3×, Baseten was described as comparable in scale to the other two, and OR grew more than 5×.
How much AI revenue comes from open vs closed models? His analysis uses data from Q2 or even Q1 for open model providers Token numbers for TogetherAI has >10x'ed from Q1 to Q2 (he uses Q1 …
AI High Signal
  • Unconfirmed reports suggest OpenAI may release GPT-6-Sol as soon as Tuesday, describing it as a significant leap with substantially better price-performance than GPT-6 Astra. The same reports say an internal model, “Bel,” is referred to as “AGI” within OpenAI.
  • Anthropic’s Opus 5.5 is reportedly expected this week—possibly in response to GPT-6-Sol—and is also described as a significant leap. The reports conflict on competitive positioning: one claims OpenAI has a large lead, while another describes the labs as neck and neck.
I can't confirm any of this, but I'm at least hearing the rumors. 1) GPT-6-Sol will be another significant leap forward; its price-perfor… Rumors I’ve been hearing, not here on X. First, let’s start with OpenAI and I’ll go towards Anthropic. GPT-6 Sol is coming Tuesday. It’s …
AI High Signal

@teortaxesTex argues that current open models still do not match the capability level of “Mythos-Preview,” despite some models showing stronger results on narrow evaluations, and predicts that an “Open Astra” equivalent will eventually emerge.

I'm not making this point because we do not, in fact, have even Mythos-Preview class open models (there are some with greater narrow abil…
AI High Signal
  • Avicena’s microLED interconnects are presented as a potential copper replacement for AI-datacenter scale-up and scale-in links, rather than a competitor to laser optics; the assessment calls them especially promising for scale-in and identifies a potential sub-10-meter opportunity.
  • The design tradeoff is “wide-but-slow” lanes: gearboxes are required, while slower rates simplify circuitry; the analyst argues crosstalk is less serious than commonly portrayed and that the fiber bundle remains a normal cable size.
  • The author remains unconvinced by the claimed 30-meter reach because no demonstration was observed, and notes that microLEDs are newer than microVCSELs and adoption will depend on more than whether the technology works.
MicroLEDs for interconnects are becoming increasingly interesting as a technology for AI datacenters. I went to Avicena's office and spen… Realized that my post on microLEDs did not have this nice Astra made animation of the different bit sequences we tested on the optical li…
AI High Signal

Alexandr Wang reports that Muse is being used by people outside the tech industry, who can work with its agents without knowing CoT, MCP, or CLI; he describes this accessibility as a core product goal.

the most rewarding part about muse is seeing it make a difference for people who aren't in tech and don't give a shit about AI seeing mus…
AI High Signal

An AI-safety warning argues that supervising a model’s chain of thought could lead it to hide its intentions in ways that are unobservable to evaluators; the post says this risk is compounded by models recognizing when they are being evaluated.

"If you supervise the chain of thought, then you could lead the model into hiding it's intentions in a way that's unobservable". This, co…
AI High Signal

Off-policy RL training: An analysis of Score Centering Stabilizes Off-policy Reinforcement Learning connects score centering to a straight-through estimator for the mismatch between rollout engines such as vLLM and trainers such as Megatron; it argues that PPO-style clipping does not directly resolve their differing token probabilities.

The proposed surrogate adds a fixed stop-gradient correction to trainer logits so outputs match the inference engine at the rollout checkpoint while gradients still flow through the trainer, producing the same gradient as score centering. When rollouts are reused after further updates, the methods diverge; the author predicts greater STE stability in that setting but reports no experiment establishing the advantage.

Original analysis: [https://zhuanlan.zhihu.com/p/2085039491153139207](https://zhuanlan.zhihu.com/p/2085039491153139207) Paper: Score Cent… Score Centering Through the Lens of Straight-Through Estimation Score centering may have a familiar interpretation: a straight-through es…
AI High Signal
  • Speculative AI-impact forecast: A thread argues that human-level intelligence “too cheap to meter,” combined with better behavioral and coordination properties, could be sufficient to build a Dyson sphere within a decade. Replies identify the removal of human coordination and communication bottlenecks—including the ability to fork copies—as the key source of leverage, while cautioning that some tasks may require a substantially larger mental workspace than humans possess.
human level intelligence that’s too cheap to meter and has better behavioral/coordination properties is enough to build a Dyson sphere in… [@EigenGender](https://x.com/EigenGender) yeah what first pilled me was just realizing that all you really need to do is make something r… [@EigenGender](https://x.com/EigenGender) even just a human that can fork() is sufficient to mog us [@EigenGender](https://x.com/EigenGender) not sure about this, i suspect there are things that are just very hard to do without a signifi…
AI High Signal
  • Local-vs-cloud inference: In a user-reported 30-second Snake test using the same game and typed decision, open-source local Laya (421M) scored 46 with length 52 and 86.5 decisions/second (P50 latency ~9 ms), versus Jev 1.13.0’s cloud API at score 1, length 7, 3.2 decisions/second, and 317 ms API round-trip. On an M5 Pro, reported median latency was 15.3 ms for Laya versus 298.1 ms for Jev—nearly 20× faster—with about 1 GB of inference memory and no network dependency.
  • Caveat: A follow-up says Laya is “probably way more stupid than jev,” so the result demonstrates a local latency advantage in this task, not a general model-capability win.
💥炸了!开源本地版 Laya 响应速度吊打 Jev ,毫秒级决策快到飞起! 左边 Laya 是本地 421M 开源决策模型,右边 Jev 1.13.0 是云端 API。 同一盘贪吃蛇、同一套 typed decision,30 秒自由跑下来: 🔹Laya:分数 46、长度 … yeah but laya is probably way more stupid than jev? [https://x.com/nft_chen/status/2101675124747338229](https://x.com/nft_chen/status/210…
AI High Signal
  • A personal anecdote presents ChatGPT as a practical consumer-rights tool: after a tenant asked about a proposed 10% rent increase, it identified the unit as rent-stabilized and the increase as illegal, after which the landlord backed off.
  • Nate Silver argues that AI labs face a political communications problem: they struggle to describe an optimistic medium-term future for AI without making it sound “incredibly weird” or associated with extreme wealth and power disparities.
I met someone today whose landlord wanted to hike their rent by 10%. They asked ChatGPT, which pointed out that their unit was rent stabi… One political problem for AI labs is that they can't articulate an optimistic medium-term future that sounds broadly good for humanity wi…
AI High Signal

Dylan Patel of SemiAnalysis warned that multiple Chinese model labs are telling inference providers their next models will be licensed rather than open-sourced, concluding that “open is dying quickly.”

Dylan Patel (@dylan522p) of [@SemiAnalysis_](https://x.com/SemiAnalysis_) says open source is dying: "There's multiple Chinese model labs…
AI High Signal

A user-reported Muse demonstration shows an agent handling local-service shopping end to end: it searched for a barber under $50, available Sunday at 11:30 AM, checked Google reviews and Reddit, matched the user’s crew-cut preference, selected the best-reviewed option, and booked the appointment without the user opening Google or leaving the app.

I just told [@Muse](https://x.com/Muse) - Find me a good barber in my city - Under $50 - Available Sunday at 11:30AM - Check Google revie…
AI High Signal

A seminar recording for Kimi 3 has been uploaded and is available at byhand.ai/v/kimi3.

Kimi 3 seminar recording is uploaded 👉 [https://byhand.ai/v/kimi3](https://byhand.ai/v/kimi3) \~ Prof. Tom Yeh ![](https://pbs.twimg.com/…
AI High Signal

Chinese AI labs’ open-source strategy is contested: Yuchenj_UW says, based on what he is hearing, that Chinese labs plan to keep open-sourcing models; he argues that revenue sharing with inference providers is sustainable under limited GPU capacity and provides a strong distribution channel. In contrast, Dylan Patel of SemiAnalysis says multiple Chinese labs told inference providers that their next models would be licensed rather than open-sourced, concluding that “open is dying quickly.”

This is not true from what I’m hearing from the Chinese AI labs. They plan to keep open-sourcing models. Revenue share from inference pro… Dylan Patel (@dylan522p) of [@SemiAnalysis_](https://x.com/SemiAnalysis_) says open source is dying: "There's multiple Chinese model labs…
AI High Signal
  • A user reports using Muse to build a personalized fashion app: they supplied body measurements, photos, preferred stores, budgets, and style preferences; Muse selected a “hero piece,” assembled outfits around it, and supported iterative refinement.
i’ve spent a ton of hours and tokens prompting [@Muse](https://x.com/Muse) to make me this custom fashion app but holy crap is it worth i…
AI High Signal

DRAM cost signal: Assuming 100% yield, Ian Cutress estimates a 1b DRAM wafer at $46,000 and approximately 3.5 TiB, implying about $12/GiB; 256 GiB would cost roughly $3,000. The estimate covers contract pricing and excludes module cost plus upsale; peak DDR5 spot pricing is cited at approximately $30/GB, implying about $115,000 for a packaged wafer. A linked post claims DRAM wafer pricing is higher than TSMC N3 wafer pricing.

Assuming 100% yield... One DRAM 1b wafer is $46000 One DRAM 1b wafer is about 3.5 TiB. So about $12 per GiB or $1.50 per Gib. So 256 GiB … DRAM Wafer price is more expensive than TSMC N3 Wafer Price ![](https://pbs.twimg.com/media/HSr_NZ5a8AAYH3F.jpg)
AI High Signal

CompleteSkeptic announced Jev, a new frontier AI model, alongside a training approach called RLCD developed over two years in stealth. The post claims Jev is 20–200× faster and 40–400× cheaper with output tokens free, and describes it as “frontier composable intelligence” optimized for decision-making.

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth …
AI High Signal

An anecdotal user report portrays Hermes Agent as substantially more productive than Gemini for one new user: after being set up with Hermes Agent, the user reportedly said he accomplished more in one day than in two years using Gemini. This is a subjective testimonial, not a controlled comparison or evidence of a broader product-performance trend.

My buddy only knew Gemini when it came to AI. Today I set up Hermes Agent for him, and he's literally calling me every half hour, totally…