ZeroNoise Logo zeronoise
Post
DeepSeek’s Reported Cost Shock Meets a Fragmented AI Stack
21 hours ago
5 min read
2665 docs
The clearest signal is a reported 105× cost advantage for DeepSeek V4 Flash, arriving as frontier labs diversify their compute and take divergent platform positions. The brief also tracks Simile’s rapid financing, new technical teams, and the growing importance of verification and infrastructure execution.

1. Funding & Deals

Simile raised $200 million in an unusually fast, insider-led round, taking total funding to $300 million in roughly six months. Founder Jun Sung Park said the company had raised $100 million about five months earlier, was not running a process, and was preempted after insiders saw unusual traction, technical progress, and a need for more compute. Green Oaks joined after tracking the market and moved within days; Park identified Index’s Shardul as the prior-round lead and Mike Volpi and Astar among the seed backers.

The thesis is not a conventional frontier language model: Simile describes a foundation model of human behavior for simulating individuals, subpopulations, and eventually markets. It wants models that reproduce human mistakes, biases, values, and preferences, using transaction and observational data alongside randomized trials and A/B tests to model causal mechanisms and counterfactuals. Park says published validation predicted behavior and attitudes 85% as accurately as people reproduce their own, while enterprise customers have closed in about three months and used Simile to reproduce findings from three-to-six-month studies in two minutes. The diligence question is whether that data-and-causal-model loop can cross the company’s own proof-of-concept-to-production chasm.

2. Emerging Teams

Suhail’s new venture is showing both frontier ambition and infrastructure fragility. The build log records a completed seed round, validation of a basic RLVR post-training stack, a first hire, and a search for a second hire in post-training or low-level model optimization. The project began with two 8xB200 systems and later acquired 64 B300s; in the latest update, Suhail said a key research component was working but that all GPUs had been lost and scaling was delayed by networking problems. For an investor, systems reliability and access to usable compute are part of the research execution risk, not merely an operational footnote.

itnetic is a sharp early security-infrastructure wedge from a one-person team. A Czech solo developer built a Rust reverse proxy for low-rate, human-like Layer-7 attacks that evaded ordinary volume-based defenses, combining JA4+ TLS fingerprinting, half-space-tree/EWMA anomaly detection, and CDN caching. The product is live with a free tier and has handled an attack of 100,000 requests per second. The signal is not revenue yet; it is a narrowly defined operational pain, a technically differentiated implementation, and evidence of deployment under real attack conditions.

3. AI & Tech Breakthroughs

OpenAI says an internal version of Astra produced ten results on long-standing problems in mathematics and theoretical computer science. The company lists advances spanning sphere packing, coding theory, group theory, quantum games, lattice cryptography, Ramsey numbers, and extremal graph theory. It says the total discovery-token cost would have been roughly $2,000 at Sol API rates; humans prepared the manuscripts, and the model formalized each argument in Lean certificates. This is a significant capability signal, but the investable question is whether independent mathematicians can reproduce and extend the work: OpenAI itself says it takes responsibility for correctness while asking the mathematical community to engage with the results.

DeepSeek V4 Flash is turning the cost-performance story into a deployment story, though the headline remains contested. An analysis cited by @kimmonismus reports that it completes the same benchmark tasks as Fable 5 at 105× lower total cost; Perplexity CEO Aravind Srinivas called two-orders-of-magnitude improvements rare and significant. A community post lists $0.09 input and $0.18 output per million tokens with a one-million-token context, while community recipes report serving the 284-billion-parameter model on one DGX Spark at 1,000 tok/s prefill and 59 tok/s in multi-agent serving; a two-Spark FP8 setup reports 82 tok/s single-stream. The counter-signal matters: Bindu Reddy calls the model “benchmark maxxed,” and another commenter rejects the comparison with Opus 4.8. Treat the 105× figure as a reported benchmark-cost result, not yet as settled capability equivalence.

4. Market Signals

The frontier compute stack is diversifying away from Nvidia. Nathan Benaich’s refreshed State of AI compute index, with a cutoff of August 1, says Anthropic added up to 2 GW of AMD MI450s, making non-Nvidia silicon 7 of its 8 GW of contracted compute; it puts OpenAI’s non-Nvidia share at 16.75 of 26.75 GW across AMD, Broadcom, and Cerebras. The implication is not that Nvidia has been displaced, but that accelerator mix, software compatibility, and supply access are becoming first-order diligence variables for model and infrastructure companies.

Platform strategy is splitting between “intelligence as a utility” and vertical integration. Garry Tan describes OpenAI’s current direction as an open platform offering intelligence on tap, while a separate post says Anthropic has been telling CEOs, VCs, and startups that it does not see the model and the application or harness as separate companies—and will therefore compete with its customers. These are operator interpretations rather than formal strategy documents, but they give application investors a concrete set of questions: how portable is the product across models, and when does the model vendor become the most dangerous competitor?

AI adoption may require a longer learning horizon than the financing cycle. Exponential View models three companies with the same starting economics and a 5% hit rate but different learning practices: after two years all are losing similar amounts, the eventual loser looks best in year five, and it takes eight years to see which approach produces outsized ROI. The same issue describes a $45 billion, roughly four-times-levered AI-capex fund that was forced to liquidate after the Philadelphia Semiconductor Index fell 28.6% from its June peak, while noting that the unwind does not prove the underlying thesis wrong.

A separate investor-sentiment signal is emerging in robotics: one post says funds are rewriting 2024 humanoid theses toward vertical-specific solutions, and Bain Capital Ventures’ Ajay Agarwal endorsed it with “Yup.”

5. Worth Your Time

  • Simile interview: the causal-data thesis. Park explains why the company wants models that reproduce human behavior rather than optimize for super-rational intelligence, and why transaction data, experiments, and counterfactuals matter.
  • The agent-artifact thread. A builder says the bottleneck in AI-assisted engineering is preserving intent, specifications, provenance, review, and knowledge transfer—not another context-window increase. The proposed durable unit is an artifact with an owner, version, and acceptance test, reinforced by human approval and diff review before changes are committed.

  • Karpathy’s Opus 5 world-building experiment. With a roughly $10, one-million-token budget, Opus spent about two hours writing 5,500 lines of Three.js to render a procedural Lord of the Rings scene; the same experiment exposes the remaining weakness in multimodal self-audit, because the model could not efficiently perceive or play-test the world it created.

DeepSeek’s Reported Cost Shock Meets a Fragmented AI Stack
Research extraction

Direct answer: OpenAI's announcement claims exactly ten advances in mathematics/theoretical computer science, all attributed to an internal version of Astra, each formalized in a Lean certificate, with the discovery token cost estimated at roughly $2,000 at Sol API rates.

  • Count and scope: "Today, we are sharing a selection of ten results to problems that have been open and have seen no progress on the main result for at least a decade, and in most cases much longer."
  • Advances: (1) high-dimensional sphere packing upper bounds down to the Cohn–Elkies threshold; (2) exponentially improved bounds for binary codes at any prescribed minimum distance and analogous results for high-dimensional spherical codes; (3) construction establishing existence of non-sofic groups; (4) disproof of Connes's rigidity conjecture; (5) new lower bounds for permanent via arithmetic circuits/formulas, including a formula lower bound of order n^4/log n; (6) exponential parallel repetition theorem for general two-player quantum games; (7) polynomial-factor hardness of approximation for the closest vector problem; (8) determining in every dimension the maximum volume of a convex body whose centroid is its only interior lattice point (Ehrhart's volume conjecture); (9) superexponential lower bound for multicolor triangle Ramsey numbers, resolving Erdős problem 183; (10) results on compactness and degeneracy conjectures in extremal graph theory, resolving Erdős problems 146 and 180.
  • Attribution to Astra: The results "were achieved by an internal version of Astra, our next major model."
  • Formal verification: "Afterward, the model formalized each argument in a Lean certificate," and OpenAI releases each solution's model narration of its thinking process. OpenAI adds it "helped prepare the manuscripts and formalize the proofs in Lean" and takes responsibility for correctness, while the mathematical arguments themselves were generated by its system.
  • Inference cost: "The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates."
  • Uncertainty/gaps: The bundle contains only this self-reported announcement; no independent verification, per-problem cost breakdown, or external replication is included.
Ten advances in mathematics and theoretical computer science | OpenAI
20VC with Harry Stebbings
  • Simile, the human-behavior simulation startup, raised a $200M round in a few days without running a formal process, bringing total funding to $300M in ~6 months; the ~$100M round ~5 months earlier was led by Index's Shardul, the new round brought in Neil Meta's team at Green Oaks (which had been tracking the market), and seed was led by Mike Volpe's firm and Astar. Founder Jun Sung Park says he was not planning to raise but insiders preempted on traction and compute needs; seed, Series A, and the new round all closed within a year .
  • Founding team: Park left Stanford in June 2025 after leading research on agents/simulations; co-founders are Michael Bernstein (ImageNet co-author, human-centered AI leader), Percy Leung (coined the term 'foundation model'), and Lainey Allen (go-to-market). Bernstein and Leung, Park's former PhD advisors, joined him, and Park says his core research team stayed together across six years of projects .
  • Technology: Simile is building a 'foundation model of human behavior' to simulate individuals, subpopulations, and eventually entire markets, descended from the 2023 Smallville experiment - 25 GPT-3.5-driven NPCs in a game town with explicit memory, planning, and reflection who self-organized a Valentine's Day party, which Park calls one of the earliest agentic systems to make those concepts explicit. Unlike frontier LLMs (the 'CPU of intelligence'), Simile wants models that make the same mistakes and carry the same biases as humans - 'the subjective half of their brain' - trained on causal data (A/B tests, RCTs) because, in Park's view, 'no one really cares about prediction'; people want to shape the future. Published validation (end-2024) predicts behaviors/attitudes 85% as accurately as people reproduce their own, starting the synthetic-panels field, and the production model now costs ~100x less to run than before .
  • Traction and market: enterprise first (e.g., CVS's VP of Insights), with deals closing in ~3 months; on first calls Simile re-derived findings from 3-6-month studies in two minutes, and value is framed as preventing decisions that could cost hundreds of millions of dollars ('a true painkiller'). Park expects synthetic panels to surpass the human-panel market, unlocking the ~95% of experiments never run, and predicts single simulation sessions costing $10-20M to run will sell for $100M in 2-3 years; long-term vision is 'representation at scale' - a 'replicable twin' of every person enabling new kinds of policy and companies .
  • Investor signals: Park calls AI labs without a clear path to world impact 'overheated' (risk of being interesting research projects, not viable companies); underinvested areas he flags are defensible data strategies and the inference/chip hardware layer - specifically bullish on Etched as it exits stealth. He concedes parts of the market are 'quite frothy' and anchors on fundamentals, notes frontier-lab researchers now earn tens of millions (startups must compete on vision and impact), and advises investors to back academic researchers 'married to impact' rather than to a problem .
The Best AI Companies Have Unique Data Acquisition Strategies | Simile Co-founder & CEO
andrew chen
  • a16z's Andrew Chen is testing DeepSeek V4 Flash 0731 on his dual NVIDIA DGX Sparks .
  • Community recipes now run the 284B-parameter DeepSeek V4 Flash GA (0731) on DGX Spark hardware: a MiaAI Lab recipe for 2x DGX Sparks with official FP8 weights hits 82 tok/s single-stream (~95 peak) and 135 tok/s at 3 concurrent sessions ; a one-command, quantized single-node version claims ~1,000 tok/s prefill and 59 tok/s multi-agent serving on one Spark .
  • Chen's Hermes agent set up the 2x DGX Spark deployment overnight from the instruction post — an early signal of agents autonomously handling AI-infrastructure installs .
just installed DeepSeek V4 Flash 0731 on my dual DGX sparks for some weekend testing! Let’s go!! Run DeepSeek v4 Flash GA (0731) on 2x [@NVIDIAAI](https://x.com/NVIDIAAI) DGX Sparks with easy ✨ 82 tok/s single stream, peak \~95 tok/s … ⚡ If you have or want a DGX Spark, this post might be the most important one you come across this month. Serve the BEST model available (… Quantized single node version [https://x.com/bleysg/status/2083448647908569321](https://x.com/bleysg/status/2083448647908569321) Looks like there’s a few single node Spark setups now with aggressive pruning and quantization so that’s cool too Just gave Hermes agent this post and it set it up overnight with no issues: [https://x.com/miaai_lab/status/2083326111144906836](https://…
Suhail
  • Former Mixpanel CEO/founder Suhail is starting a new AI venture: seed round secured , domain/name acquired .
  • Team: first hire made; hiring #2 for post-training (RLVR/OPSD) or low-level model optimization ; earlier hunting for employee #1 .
  • Technical: basic RLVR post-training stack validated ; an 'autonomous AI scientist' is being pointed at new optimizations ; key research works and needs scaling, but all GPUs were lost ('GPU poor') and more compute is delayed by networking issues .
  • Compute: started with 2 8xB200s and later acquired 64 B300s .
  • Founder background: previously worked on image models, now reviewing fundamentals .
5/ Funding secured. Seed round done. 6/ domain / name acquired 9/ made the first hire ❤️ Looking for [#2](https://x.com/hashtag/2): post training (RLVR/OPSD/etc) or low level model optimization 4/ On the hunt for employee [#1](https://x.com/hashtag/1) :) 8/ basic RLVR post-training stack validated ![](https://pbs.twimg.com/media/HL6LiPhboAIl8wK.jpg) 3/ Time to let my autonomous ai scientist rip on some new optimizations ![](https://pbs.twimg.com/media/HKZJymUa0AAp-8r.png) 12/ got a very key piece of research working and need to scale it up; lost all my GPUs today though so now I am GPU poor more coming but … 1/ it all started w 2 8xB200s excited to be back in the game again 10/ 64 B300s acquired - if you search hard enough, you'll find what you need 2/ spent a lot of time reviewing the absolute fundamentals again; missed a lot in the world while working on image models there’s so much…
andrew chen

Andrew Chen (a16z-affiliated) says users are increasingly asking an LLM for the exact right answer instead of clicking through 10 different blue links cluttered with popups, ads, and flashing text, and welcomes the end of the 'spammy blue links paradigm,' especially for travel, tech support, and movie reviews .

Isn’t it funny we used to look up stuff and click on 10 different blue links, all overwhelmed with popups, ads, and flashing text, instea…
Bindu Reddy

Bindu Reddy, CEO of AI agent/supercomputer builder Abacus AI, urges OpenAI to quickly ship Astra, its 'Fable class' model, warning that Fable adoption is growing rapidly and users will be hard to cut over or change if OpenAI waits too long . She also claims a self-improving agent on Fable 5 'can literally solve any problem already' . The post comes from a competitor building rival AI agent technology, so the adoption and capability claims are self-interested and unverified.

OpenAI better drop Astra, their Fable class model quickly.... Fable adoption is growing rapidly, and it will be hard for users to cut ove…
Paul Graham

Paul Graham notes a recent YC startup spent $250k on a domain, which was 1/26 of the $6.5m they raised after YC; he calls it a sign of the current era where one shocking number counterbalances the other .

A recent YC startup spent 250k on a domain. I was slightly shocked. But when I asked how much they raised after YC, the answer was 6.5m. …
Garry Tan

Garry Tan flags a 2026 "vibe shift": OpenAI is positioning as the open platform, with AI as "intelligence on tap as a utility," contrasted against signaling that full-stack integration is optimal . In a post he quotes, @chetanp says Anthropic has been broadcasting publicly and privately to CEOs, VCs, and startups that they don't see a world where the application (harness, etc.) and the model are separate companies — implying Anthropic will compete with its own customers . Together these signals point to a strategic divergence: OpenAI as an open/utility layer versus Anthropic pursuing vertical integration, with direct implications for startups building application layers atop model APIs and for investors assessing platform risk.

Most interesting 2026 vibe shift is OpenAI actually looking to be the open platform Note the marked difference: intelligence on tap as a … It's subtlety wrapped in technical terms that Anthropic has been broadcasting both publicly and privately to CEOs, VCs, startups. They do…
andrew chen

Andrew Chen (a16z) argues that as the cost of creating something drops toward zero, the cost shifts elsewhere — toward skilled verification of proofs, code review, human attention for videos, and courts/judges/lawyers for lawsuits — implying value concentrates in these scarce human bottlenecks .

Unlimited mathematical proofs but limited skilled mathematicians who can verify them Unlimited code but limited skilled programmers to re…
sarah guo

Sarah Guo (@saranormous) frames BCI as 'increasingly backable engineering,' pulled forward by AI — models we interact with, translation architectures, and predictors of how perturbations reshape experience — noting that all experience of reality is mediated by brain electrical activity . In a quoted earlier post, she says BCI and new hardware (silent speech, glasses, ultrasound devices, deeper research) are 'suddenly back in vogue' and predicts we will eventually 'silently think at (with?) the AI' . Echoing Max Hodak, she argues the space needs companies to generate hundreds of millions in revenue to 'open the floodgates of investor capital' for long-term goals .

Our entire experience of reality is mediated by electrical activity in the brain. I’m inspired by BCI becoming increasingly, backable eng… BCI and new hardware is suddenly back in vogue (silent speech, glasses, ultrasound devices, deeper research) but it should be. We will ab… as [@maxhodak_](https://x.com/maxhodak_) has said, we need companies in the space to generate hundreds of millions of revenue, and that w…
@jason

Jason Calacanis (@Jason) predicts "Every product you ever dreamed about is gonna be built in the next 12 months" , asserting that software abundance already exists and hardware abundance will arrive in less than 12 months —a directional market signal for early-stage investing.

Every product you ever dreamed about is gonna be built in the next 12 months Software abundance is already hwtr, hardware abundance in le…
sarah guo
  • Karpathy reports giving Opus 5 the first paragraph of Lord of the Rings with a 1M-token budget (~$10); Opus spent ~2 hours writing 5,500 lines of Three.js code that procedurally rendered the story — janky but functional.
  • He frames this as a shift from "no one would ever do this" to "sure, why not, it's ~free," and points toward hyper-custom, on-demand worlds — "an ephemeral GTA of X on demand" — where users could join as spectator NPCs or characters.
  • He flags a key capability gap: LLMs can't efficiently audit their own work in world/game domains because they can't natively perceive videos or play games; Opus 5 had to slowly take screenshots and produced jank.
  • Sarah Guo replied to Karpathy: "it's time for The Mind Game".
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize … [@karpathy](https://x.com/karpathy) it’s time for The Mind Game
Leo Polovets

Leo Polovets (GP @HumbaVC) defends SF tech as a deep meritocracy — "flawed, but WAY better than anywhere else," with many successful people starting with nothing . He argues that raising tens or hundreds of millions in one's early twenties reflects an ecosystem putting "that much faith into someone's ability and demonstrated execution" . He distinguishes earned pedigree (top school via merit; MIT has 4% acceptance) from unearned billionaire-family birth (~1 in 100k) . Quoted in the thread, @Celesteamadon observes that SF founders often pour 100% of their lives into work, then "come up for air in their 30s" without friend groups or dating experience .

This clip is getting a lot of flak. "SF isn't a meritocracy LOL" "pedigree is everything" I disagree. SF's meritocracy is like democracy:… The fact that you can raise tens or hundreds of millions in your \*early twenties\* is a case in point. Where else will an ecosystem put … Sure, some founders come from MIT/etc, but that's very different from "born into a billionaire family" pedigree that's not merit-based. Y… "SF is a deep meritocracy. people don't really care how much money you have or care where you went to school" - [@Celesteamadon](https://…
@jason

Jason Calacanis (@Jason) states that the commodification of intelligence means the democratization of intelligence — a high-level thesis signal from a prominent startup investor.

Commodification of intelligence means democratization of intelligence
Ajay Agarwal

Robotics VCs are quietly pivoting investment theses from humanoid warehouse robots to vertical-specific solutions after funding 'vaporware demo bots' in 2024; @antopatrex1 predicts LP letters will reveal the shift, and @ajay_bcv (Bain Capital Ventures) endorsed the post ('Yup') . This is a cautionary signal for robotics investing: the same partners/funds are repositioning from 'humanoids will replace warehouse workers by 2026' to 'we always believed in vertical specific solutions' .

every robotics VC that funded vaporware demo bots in 2024 is quietly rewriting their investment thesis right now. same partners, same fun… Yup [https://x.com/antopatrex1/status/2083334392227709287](https://x.com/antopatrex1/status/2083334392227709287)
Exponential View
  • Leopold Aschenbrenner's AI-capex fund Situational Awareness LP — started late 2024 by an AI researcher with no hedge fund experience — peaked at $45B and was forced to liquidate this week after betting on the AGI build-out at roughly four times leverage (thesis: superintelligence would put trillions into compute, chips and power) . The unwind followed the Philadelphia Semiconductor Index falling 28.6% from its June peak as software gained ; the source notes the liquidation is not proof the thesis is wrong .
  • AI adoption model: three types of AI-adopting companies with the same starting economics and 5% hit rate but different learning practices — all lose similar money at two years, the eventual loser looks best at year five, and it takes eight years to identify the approach with outsized ROI . The post frames Zuckerberg as the "king of the side quest": Meta doesn't need each project to succeed as long as experiments deepen its infrastructure and inform the next move; system builders that compound learnings over time have the winning formula .
🔮 Leopold & exponential markets; transformative GLP-1s; runaway AI & the future of safety++ #595
martin_casado

Martin Casado (a16z) observes that AI model capabilities/releases appear to be accelerating, but when divided by the money going into the labs, progress looks sublinear; organizations get less efficient at scale, so the overall picture is not an obvious "take off" .

It appears model capabilities / releases are accelerating. But if you divide by the money going into the labs, it looks sublinear. Clearl…
@jason

Jason Calacanis (@Jason) posted that he sees no upper limit for on-demand intelligence, a bullish investor sentiment signal for AI-driven services.

I don't see an upper limit for on-demand intelligence
@jason

VC Jason Calacanis (@Jason) expressed excitement about the recent pace of AI model releases, noting they are "good, cheap and fast" and quipping that no one will sleep if they keep dropping at this rate, reflecting strong investor enthusiasm for rapid AI progress .

No one is gonna sleep again if these models keep dropping this good, cheap and fast! What did you lunatics do last night?!
The community for ventures designed to scale rapidly | Read our rules before posting ❤️
  • Fintech startups generally aren't valued at SaaS multiples — valuations often track AUM or transaction volume — and they usually hit headwinds when raising Series B–D; failure to raise leads to staff trims, shutdown, or fire sale .
  • A warning that most fintechs are "basically all just marketing companies except for the rare few": credit cards, ledger/APR, and ACH are commoditized API services (e.g., Fiserv), so growth reduces to marketing dollars for customers/deposits to capture the APR spread, and doing it fast enough is the hard part — incumbents quickly copy anything that affects their business .
  • Cautionary data point: one commenter joined a fintech at a $400M valuation in 2019, saw it peak at $2B in 2022, and the company later sold for roughly $20M .
  • Early-stage companies treat profitability as a mirage — any brief profitability is followed by more spend to keep growth at a pace that satisfies current and future investors .
  • Context: the thread's subject is a post-Series A fintech (~$100M valuation) building B2B SaaS for banks/businesses to execute large-quantity trades (no retail customers), targeting Series B but ultimately aiming for an acquisition .
I mean that fintechs aren’t generally valued at saas multiples. Their valuation depends on a few factors but often times depends on AUM o… Unless the fintech is doing something truly transformative, they are basically all just marketing companies except for the rare few. Ever… Yep. It’s extremely challenging to get customers quickly enough. Other established large financials will quickly copy anything you do if … Joined fintech myself at $400m in 2019, peaked at 2b 2022 and sold for like 20m Profitability is a mirage at an early stage company. It’s either just on the horizon or you reach it for a minute and then need to spend … Leaving FAANG for Fintech Startup Post Series A (I Will Not Promote) Theyre in the trading space for banks and business to execute trading on their SaaS. No retail customers, focused on large quantity trades.