ZeroNoise Logo zeronoise
Post
DeepSeek’s Reported Cost Shock Meets a Fragmented AI Stack
17 hours ago
5 min read
2665 docs
The clearest signal is a reported 105× cost advantage for DeepSeek V4 Flash, arriving as frontier labs diversify their compute and take divergent platform positions. The brief also tracks Simile’s rapid financing, new technical teams, and the growing importance of verification and infrastructure execution.

1. Funding & Deals

Simile raised $200 million in an unusually fast, insider-led round, taking total funding to $300 million in roughly six months. Founder Jun Sung Park said the company had raised $100 million about five months earlier, was not running a process, and was preempted after insiders saw unusual traction, technical progress, and a need for more compute. Green Oaks joined after tracking the market and moved within days; Park identified Index’s Shardul as the prior-round lead and Mike Volpi and Astar among the seed backers.

The thesis is not a conventional frontier language model: Simile describes a foundation model of human behavior for simulating individuals, subpopulations, and eventually markets. It wants models that reproduce human mistakes, biases, values, and preferences, using transaction and observational data alongside randomized trials and A/B tests to model causal mechanisms and counterfactuals. Park says published validation predicted behavior and attitudes 85% as accurately as people reproduce their own, while enterprise customers have closed in about three months and used Simile to reproduce findings from three-to-six-month studies in two minutes. The diligence question is whether that data-and-causal-model loop can cross the company’s own proof-of-concept-to-production chasm.

2. Emerging Teams

Suhail’s new venture is showing both frontier ambition and infrastructure fragility. The build log records a completed seed round, validation of a basic RLVR post-training stack, a first hire, and a search for a second hire in post-training or low-level model optimization. The project began with two 8xB200 systems and later acquired 64 B300s; in the latest update, Suhail said a key research component was working but that all GPUs had been lost and scaling was delayed by networking problems. For an investor, systems reliability and access to usable compute are part of the research execution risk, not merely an operational footnote.

itnetic is a sharp early security-infrastructure wedge from a one-person team. A Czech solo developer built a Rust reverse proxy for low-rate, human-like Layer-7 attacks that evaded ordinary volume-based defenses, combining JA4+ TLS fingerprinting, half-space-tree/EWMA anomaly detection, and CDN caching. The product is live with a free tier and has handled an attack of 100,000 requests per second. The signal is not revenue yet; it is a narrowly defined operational pain, a technically differentiated implementation, and evidence of deployment under real attack conditions.

3. AI & Tech Breakthroughs

OpenAI says an internal version of Astra produced ten results on long-standing problems in mathematics and theoretical computer science. The company lists advances spanning sphere packing, coding theory, group theory, quantum games, lattice cryptography, Ramsey numbers, and extremal graph theory. It says the total discovery-token cost would have been roughly $2,000 at Sol API rates; humans prepared the manuscripts, and the model formalized each argument in Lean certificates. This is a significant capability signal, but the investable question is whether independent mathematicians can reproduce and extend the work: OpenAI itself says it takes responsibility for correctness while asking the mathematical community to engage with the results.

DeepSeek V4 Flash is turning the cost-performance story into a deployment story, though the headline remains contested. An analysis cited by @kimmonismus reports that it completes the same benchmark tasks as Fable 5 at 105× lower total cost; Perplexity CEO Aravind Srinivas called two-orders-of-magnitude improvements rare and significant. A community post lists $0.09 input and $0.18 output per million tokens with a one-million-token context, while community recipes report serving the 284-billion-parameter model on one DGX Spark at 1,000 tok/s prefill and 59 tok/s in multi-agent serving; a two-Spark FP8 setup reports 82 tok/s single-stream. The counter-signal matters: Bindu Reddy calls the model “benchmark maxxed,” and another commenter rejects the comparison with Opus 4.8. Treat the 105× figure as a reported benchmark-cost result, not yet as settled capability equivalence.

4. Market Signals

The frontier compute stack is diversifying away from Nvidia. Nathan Benaich’s refreshed State of AI compute index, with a cutoff of August 1, says Anthropic added up to 2 GW of AMD MI450s, making non-Nvidia silicon 7 of its 8 GW of contracted compute; it puts OpenAI’s non-Nvidia share at 16.75 of 26.75 GW across AMD, Broadcom, and Cerebras. The implication is not that Nvidia has been displaced, but that accelerator mix, software compatibility, and supply access are becoming first-order diligence variables for model and infrastructure companies.

Platform strategy is splitting between “intelligence as a utility” and vertical integration. Garry Tan describes OpenAI’s current direction as an open platform offering intelligence on tap, while a separate post says Anthropic has been telling CEOs, VCs, and startups that it does not see the model and the application or harness as separate companies—and will therefore compete with its customers. These are operator interpretations rather than formal strategy documents, but they give application investors a concrete set of questions: how portable is the product across models, and when does the model vendor become the most dangerous competitor?

AI adoption may require a longer learning horizon than the financing cycle. Exponential View models three companies with the same starting economics and a 5% hit rate but different learning practices: after two years all are losing similar amounts, the eventual loser looks best in year five, and it takes eight years to see which approach produces outsized ROI. The same issue describes a $45 billion, roughly four-times-levered AI-capex fund that was forced to liquidate after the Philadelphia Semiconductor Index fell 28.6% from its June peak, while noting that the unwind does not prove the underlying thesis wrong.

A separate investor-sentiment signal is emerging in robotics: one post says funds are rewriting 2024 humanoid theses toward vertical-specific solutions, and Bain Capital Ventures’ Ajay Agarwal endorsed it with “Yup.”

5. Worth Your Time

  • Simile interview: the causal-data thesis. Park explains why the company wants models that reproduce human behavior rather than optimize for super-rational intelligence, and why transaction data, experiments, and counterfactuals matter.
  • The agent-artifact thread. A builder says the bottleneck in AI-assisted engineering is preserving intent, specifications, provenance, review, and knowledge transfer—not another context-window increase. The proposed durable unit is an artifact with an owner, version, and acceptance test, reinforced by human approval and diff review before changes are committed.

  • Karpathy’s Opus 5 world-building experiment. With a roughly $10, one-million-token budget, Opus spent about two hours writing 5,500 lines of Three.js to render a procedural Lord of the Rings scene; the same experiment exposes the remaining weakness in multimodal self-audit, because the model could not efficiently perceive or play-test the world it created.

DeepSeek’s Reported Cost Shock Meets a Fragmented AI Stack
Summary
Coverage start
1 day ago
Coverage end
17 hours ago
Frequency
Daily
Published
15 hours ago
Reading time
5 min
Research time
13 hrs 4 min
Documents scanned
2665
Documents used
27
Citations
33
Sources monitored
119 / 120
Insights
133
View
Skipped contexts
213
View
Source details
Source Docs Insights Status
Hunter Walk 0 0
SaaStr 2 1
andrewchen 0 0
VC Adventure 0 0
Elad Blog | Substack 0 0
AVC 0 0
Above the Crowd 0 0
Entrepreneur Ride Along 85 5
r/SideProject - A community for sharing side projects 381 25
Future(s) Studies 1163 27
Artificial Intelligence (AI) 99 6
Software As a Service Companies — The Future Of Tech Businesses 563 30
Investing In AI 0 0
Big Technology 0 0
The Gradient 0 0
Import AI 0 0
Sam Altman 0 0
The community for ventures designed to scale rapidly | Read our rules before posting ❤️ 126 6
Co-Founder: Find Your Co-Founder Here 0 0
Entrepreneur 73 0
Naval 0 0
Machine Learning 19 3
Deep Learning 16 3
Natural Language Processing 42 0
Venture capital news and articles, for the VC industry 0 0
Newcomer 0 0
Jerry Liu 0 0
Harrison Chase 0 0
Cristóbal Valenzuela 1 1
Amjad Masad 2 0
Arthur Mensch 0 0
clem 🤗 0 0
Aidan Gomez 0 0
Kanjun 🐙 0 0
Suhail 12 1
Guillaume Lample @ NeurIPS 2024 0 0
Clouded Judgement 0 0
Bindu Reddy 2 2
Parag Agrawal 0 0
Harry Stebbings 1 1
Keith Rabois 0 0
Fred Wilson 0 0
Brad Feld 3 0
Exponential View 2 1
The Pragmatic Engineer 0 0
Latent.Space 0 0
Mark Suster 0 0
Benedict Evans 0 0
Allie K. Miller 0 0
Elizabeth Yin 💛 0 0
Roelof Botha 0 0
Andrew Reed 0 0
Luciana Lixandru 0 0
The Pragmatic Engineer 0 0
Elad Gil 4 2
Nathan Benaich 4 1
sarah guo 7 3
@jason 14 5
Vinod Khosla 0 0
Daniel Gross 0 0
Ann Miura-Ko 🦖 0 0
Mike Volpi 0 0
Aravind Srinivas 2 1
Ajay Agarwal 2 1
Leo Polovets 4 1
David Sacks 0 0
Lenny's Newsletter 0 0
Interconnects 0 0
Not Boring by Packy McCormick 0 0
Marc Andreessen 🇺🇸 0 0
Chris Dixon 0 0
Sriram Krishnan 2 0
a16z 0 0
benahorowitz.eth 0 0
martin_casado 15 1
andrew chen 9 3
Scott Kupor 0 0
David Ulevitch 🇺🇸 1 0
Dalton Caldwell 0 0
Y Combinator 0 0
Jessica Livingston 0 0
Paul Graham 4 1
Invest Like The Best 0 0
Garry Tan 2 1
Michael Seibel 0 0
Sam Altman 2 0
TechCrunch 0 0
Plug and Play Tech Center 0 0
No Priors: AI, Machine Learning, Tech, & Startups 0 0
Lex Fridman 0 0
Lightspeed Venture Partners 0 0
500 Global 0 0
Google for Startups 0 0
ThisWeekinStartups 0 0
Two Minute Papers 0 0
My First Million 0 0
Lenny's Podcast 0 0
All-In Podcast 0 0
Garry Tan 0 0
Y Combinator 0 0
Acquired 0 0
Foundation Capital 0 0
20VC with Harry Stebbings 1 1
Sequoia Capital 0 0
Greylock 0 0
Stanford eCorner 0 0
a16z 0 0
Jeremy Howard 0 0
Aravind Srinivas 0 0
Cassie Kozyrkov 0 0
Andrej Karpathy 0 0
Alexandr Wang 0 0
Naval Ravikant 0 0
Clément Delangue 0 0
Elad Gil 0 0
Fei-Fei Li 0 0
Andrew Ng 0 0
Demis Hassabis 0 0
Sam Altman 0 0
Yann LeCun 0 0