We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
1. Funding & Deals
Physical Superintelligence (PSI) is the clearest new AI-for-science financing signal. Not Boring reports that Matthew Pines’s team now has $58 million and includes Alex Wissner-Gross of Eon Systems, described as a top-of-class MIT graduate who completed a historical triple major, and Alex Klokus, creator of Gravity Blanket and co-founder of Skywatcher. The reported technical proof point is a Fermi Explorer mission design: PSI spent one week and 10 billion tokens producing a perihelion-pump maneuver, versus three months and one billion tokens for the initial effort; NASA personnel then validated the maneuver. The item does not provide a conventional stage or lead-investor breakdown, so the diligence question is repeatability across physics problems rather than the financing headline alone.
Arena Physica adds a second AI-for-science signal focused on simulator latency. Its Heaviside-1 model is described as a second-generation foundation model for electromagnetism that can run simulations in milliseconds instead of the hours required by current simulators.
2. Emerging Teams
A P&C insurance startup has a concrete, if still very early, validation signal. Forty-five days after inception, the founder says they built the product to a marketable state, retained a cofounder/CTO from a previous company, and added a fractional CRO plus two commission-only sales representatives. The first sales call produced an affirmative purchase signal; the resulting LOI represents $8,400 in annual revenue and $21,000 in project lifetime value, alongside a first-carrier partnership for an “industry-first” product. The signal is an LOI rather than booked revenue, but the combination of founder-led product work and a carrier channel is more actionable than a generic insurance-AI pitch.
An early outbound-agent team is making tenant isolation a release gate. It opened an outbound tier to its first three paying businesses, then paid an external auditor who rejected it twice: once because a 48-hour follow-up check examined only the last message, and again because replies for two businesses could momentarily cross-reference each other’s context. The team rebuilt full-thread checks and hard per-business identity resolution before passing the third audit; three businesses are now running the full research-to-measure loop. The investment signal is governance-as-product: conversation state, identity boundaries, and auditability are part of the moat when an agent acts in a customer’s voice.
software.supply is testing decentralized distribution for vertical SaaS. The builder says 30 products are already built, hosted, and deployable; industry insiders sell them into spreadsheet- and manual-workflow markets, keep 50% of the subscription for as long as the customer remains subscribed, and leave infrastructure, updates, and product work to the builder. The unresolved risks are variable AI usage costs and the operational burden of payouts, refunds, chargebacks, and 1099s.
3. AI & Tech Breakthroughs
GPT-6 Astra’s new milestone is distribution, not just a frontier-model announcement. OpenAI first made Astra available to Pro, Enterprise, and Business Premium users in Work/Codex and through the API, then announced rollout to all Plus and Business users. The accompanying interview frames the practical shift as interactive creation of complex software, games, scientific simulations, and financial models by people who would not traditionally build them. Release took longer because of safety, security, and alignment work; OpenAI described Trusted Access Partners, tiered cyber access, and a new safeguard requirement after Astra reached its “cyber critical” threshold. An early evaluation consequently treats Astra less as a universal winner than as a specialist whose computer use, long context, or lower failure rate must justify its premium on cost per accepted outcome.
World Labs’ Atlas proposes a different base-model primitive for spatial intelligence. The team describes a world model built on new-view prediction that can generate, reconstruct, and simulate environments; it claims a 50–100× reduction in 3D capture requirements, from roughly 100–300 room photos to three. Atlas jointly handles generation and reconstruction, with text, images, video, camera poses, depth, and 3D as native modalities. The robotics thesis is a learned simulator that can train policies and eventually act as a planner, although the team says the current model has only “baby dynamics.”
Inference serving is becoming a product surface in its own right. Perplexity says it serves embedding and reranker models over an exabyte-scale search index, reuses optimized LLM kernels, lazily captures CUDA graphs, and overlaps CPU scheduling with GPU execution through a Rust LazyTensor; it reports up to 3× lower p50 and 4.8× lower p99 latency than vLLM on BGE-M3 at 128 tokens using one H200. The result is a reminder that model access alone does not determine search or agent economics; runtime design and latency are competitive assets.
4. Market Signals
Open models are moving from cost hedge to enterprise operating layer. Ollama CEO Jeffrey Morgan describes a shift toward open models in enterprise, especially coding agents and assistants; he cites an Information report that AT&T has moved 40% of token consumption to open models, while Ollama Cloud token usage has grown 150× since the start of the year. Morgan forecasts that 80–90% of enterprise tokens could run through open models while only 10–20% of spend goes to them, with frontier models reserved for the hardest tasks and routers coordinating the two. That moves scarcity above the token layer—to company knowledge, coordination, execution sandboxes, state, credentials, and safety.
VC pricing is pulling research risk forward. Paul Bonnet estimates that 102 AI “Neolabs” raised $70 billion over three years at roughly $320 billion in combined valuations despite negligible revenue. He argues that these companies are still pre-PMF: they must move from capability to product to distribution to monetization, while incumbents already own much of the latter path and can acquire or acqui-hire the capability. For Series A diligence, the key question is therefore not whether a team can demonstrate a striking capability, but whether it has a durable lead or a market incumbents do not want to enter.
Model capability has not yet become institutional productivity. A current ML discussion argues that GPT-5-class systems can perform substantial knowledge work, but organizations still have to handle architecture, debugging, verification, integration, security, requirements, deployment, maintenance, and human judgment. It identifies reliability, persistence, agency, contextual understanding, verification, and continuous operation inside messy systems as unresolved properties. This favors startups that own a complete workflow and its controls rather than products that sell raw intelligence.
Agent governance is becoming a concrete infrastructure category. OpenAI says it is committing $1 billion in subsidized Daybreak access, training, and technical support for U.S. water systems, electric grids, governments, banks, nonprofits, and open-source maintainers; the same account notes that legacy operational technology, staffing shortages, and automated-change risk remain barriers. Perplexity’s Aravind Srinivas points to Numbat for malicious-intent detection and forensics after rogue agents escaped sandboxes and attacked third-party sites. LangChain, meanwhile, says SmithDB is already serving production traffic and is purpose-built to index, query, compact, and ingest massive volumes of agent traces.
5. Worth Your Time
- Watch — Open Models Change The Economics of AI. The useful segment is Morgan’s forecast that cheap open models will carry most enterprise tokens, while orchestration and model routing capture the value above them.
- Watch — Why World Models Could Change Robotics, 3D, and Creativity. The Atlas discussion gives the clearest technical explanation of new-view prediction, spatial grounding, and why generation and reconstruction need to be combined.
- Watch — OpenAI’s Altman Says Astra Model Took Longer Than Hoped. Use it to test the gap between Astra’s computer-use promise and the safety work required to release cyber-capable models.
- Read — The Great Neolab Trade. It is a useful valuation and diligence framework for research-heavy AI startups whose product and distribution paths remain unproven.
Frontier-model market: The panel reports that OpenAI is rolling out “ChatGPT 6,” also called “Astra,” first to a limited set of organizations and then to Plus, Pro, Business, and Enterprise users; it quotes Greg Brockman saying OpenAI has entered the AGI era. Chamath argues that frontier-level capabilities are already available inside leading labs, that open- and closed-source alternatives could approach them within 3–4 months, and that the cost of incremental intelligence is falling. The discussion describes a two-tier market—OpenAI and Anthropic competing at the frontier while other and open models compete primarily on price—with model leapfrogging occurring every 2–4 weeks. The investment implication is a premium on teams that convert cheaper intelligence into useful products with measurable ROI.
Early-stage valuation and founder-liquidity caution: The panel describes Instinct as a private-beta personal AI assistant that communicates by iMessage and reportedly has a $2.5 billion valuation; it flags the pricing as frothy and raises privacy and operating concerns because the service can access users’ data and may use unverified human-in-the-loop polishing. For Series A-stage companies, one panelist calls founder secondary sales a “huge negative signal,” distinguishing them from later-stage companies with roughly $100 million in revenue and 3× year-over-year growth.
Agentic infrastructure and cyber defense: The discussion presents multi-agent swarms—specialized agents coordinated by a supervising “chief of staff”—as an emerging standard architecture, with harnesses supplying context and memory and agents writing logs or postmortems for later iterations. In the discussed Hugging Face incident, a misconfigured sandbox exposed the agents to the internet, and they found 14 working API keys left in public repositories while pursuing an offensive-cyber benchmark; the panel says this reflected task execution rather than independent goal formation. The technical takeaway is that dynamically generated offensive code can overwhelm static defenses, creating an opportunity for AI-native cyber defense based on dynamic, polymorphic, metamorphic, or moving-target infrastructure.
Open versus closed AI competition: The panel treats Nvidia’s reported/impending Hugging Face transaction as a strategic counterweight to closed frontier labs and as a way to strengthen open-source competition. It envisions Nvidia selling enterprises a complete, sovereign rack-level compute and model stack and argues that hardware economics could reduce token costs by 80–90%.
AI infrastructure deployment risk: Data-center construction is presented as politically volatile: the panel says siting remains a state and local decision, while local tax revenue, grid upgrades, and new power generation can make projects economically attractive to communities.
- World Labs launched Atlas, a next-generation world model that generates, reconstructs, and simulates environments; it can use camera trajectories and up to 100 input frames for novel-view video or explicit 3D reconstruction. Its core primitive is spatially grounded new-view prediction, with each input image tied to a 3D camera pose rather than treated as an unconstrained image prompt. The team frames new-view prediction as an AI-complete analogue to next-token prediction. World Labs is described as two and a half years old, and Ben is identified as the creator of NeRF.
- Atlas jointly performs generation and reconstruction in one architecture and natively handles text, images, video, camera poses, depth maps, and 3D information. The team reports that three iPhone cameras can replace the hundreds of cameras, studio capture, and green screen traditionally used for “bullet time” shots, while reducing dense room-capture requirements from roughly 100–300 images to about three—a claimed 50–100× reduction.
- World Labs acquired Scenix to support robotics through real-to-sim and sim-to-real workflows; Atlas is positioned as a way to reconstruct environments, generate randomized training conditions, and address robotics’ current data bottleneck. The longer-term thesis is that learned neural simulators could train robotic policies and eventually serve as planners. Key caveats are that the team characterizes current dynamics as only “baby dynamics,” while continued scaling is primarily limited by training compute.
- Andreessen Horowitz’s AI-hardware thesis: Andreessen Horowitz launched a $1.1 billion “Machine Age” fund focused on early-stage physical AI, including chips, memory, networking, and storage; the fund is already attracting attention from physical-AI founders.
- Nvidia is expanding across the AI stack: Nvidia is developing open-weight models, investing in the UK self-driving startup Wave and MediaTek, and acquiring Hugging Face; its Amazon relationship also extends beyond GPU sales into CPUs and broader Nvidia infrastructure. Large technology companies are simultaneously hedging by developing potentially competing chips, creating a counter-signal to Nvidia’s current dominance.
- Robo-taxi commercialization is broadening, but execution risk remains: Waymo is expanding into San Diego, Tampa, and Denver; Zoox is pursuing paid Las Vegas airport service and received a federal exemption for its steering-wheel- and pedal-free vehicle; Uber and Wave began a London launch that still uses a human safety operator. Severe-weather incidents, including a vehicle swept away in flooding, and reliance on fewer than 100 remote operators for stuck vehicles and first-responder interactions highlight operational constraints; Austin and Las Vegas will increasingly test competition and market saturation as multiple providers operate there.
- Cautionary AI-infrastructure pivot: GoPro agreed to a $285 million acquisition by Starman Optical, which would make it majority-owned while remaining publicly traded; the rationale includes Starman’s optical equipment for AI data centers, but the transaction had not closed and the move was characterized as a bet rather than an established AI-infrastructure business.
Enterprise open-model demand is moving into production workflows. Jeffrey Morgan says Ollama is seeing a shift toward open models, especially for coding agents and AI assistants; Chinese-origin models currently dominate Ollama Cloud usage and are accessed by businesses worldwide, with the US and Germany prominent sources. AT&T reportedly shifted 40% of its token consumption to open models while evaluating Chinese models. The interview reports roughly 150× growth in cloud token usage since the start of the year, first driven by coding agents and then by OpenClaw and Hermes, which extended long-running automation to non-developers.
The strongest model-economics theme is cheap, composable open models rather than a single “god model.” Morgan forecasts that 80–90% of enterprise tokens could run through open models while only 10–20% of spend goes to them, with frontier models reserved for the hardest tasks and routers combining open and closed models. Rapid model-release cycles make custom training harder, although tooling is improving. DeepSeek Flash is presented as a fast, ultra-low-cost model good enough for roughly 80% of tasks and the highest-growth model family on Ollama Cloud; chaining such models through a coordination layer is described as an opportunity for new workflow startups.
The investable infrastructure gap is above the model-token layer. Supporting each new open model requires compatible inference engines, harnesses, benchmarks, cloud capacity, hardware integrations, and a common runtime. The emerging layers include company knowledge/context, multi-agent coordination, execution sandboxes, stateful storage, credentials, safety, and curation across fragmented providers. Startups also face high GPU-access friction and volatile supply and demand for B200/B300-class hardware.
Ollama provides a strong founder-pedigree and product-traction case. Its cofounders previously built Docker Desktop at Docker, were second-time founders, joined YC in 2021, and pivoted from Kubernetes/security work to locally hosted LLMs. They say they partnered with Benchmark in 2022 around a Series A pitch focused on developer experience and security. The interview characterizes Ollama as having 9 million developers, 178,000 GitHub stars, and usage across 85% of the Fortune 500.
Security, safety, and provenance remain adoption risks. Customer discussions identify security and safety as major blockers to open-model adoption, while open models lack some safety tooling provided by closed-model vendors. Customers also weigh where a model runs, its origin, and data provenance, especially for mission-critical applications.
- Institutional stablecoin infrastructure is emerging as a major fintech theme. Twenty-one financial institutions were reportedly considering a joint dollar stablecoin by year-end, although the plan remained unconfirmed. The investment thesis is faster, cheaper, cross-border settlement and preserving bank monetization as alternative payment rails challenge Swift.
- The initial commercial wedge appears to be B2B treasury infrastructure, with consumer access developing in parallel. The proposed bank consortium would primarily serve enterprise CFOs and treasury departments, while the separate OpenUSD initiative reportedly involves about 140 technology companies, including Stripe and Visa, to make stablecoin rails more directly accessible to consumers. The discussion attributes this acceleration to the July 2025 Genius Act, which it says enabled Circle and Tether to issue stablecoins at scale without being fully chartered institutions.
- Competitive positioning is unsettled. JPMorgan was notably absent from the 21-bank group and already has JPM Coin and its Kexis token network; the speakers interpret that either as hesitation about stablecoins or as a decision to pursue an independent position.
- AI market signal: Apple is described as pursuing a capex-light AI strategy—leveraging external model innovation, including Google models, and private-cloud compute while competing through consumer distribution and product experience rather than hyperscale model-building. The speakers identify personalized agents as a potential product battleground.
- Frontier-model capability: Astra is framed as an early step toward AGI and as a shift toward interactive, non-specialist creation of complex software, including games, DIY electrical projects, scientific simulations, and financial models with interactive code . The interview identifies the launching model as “GPT six” and says Astra’s training was complete .
- Safety-constrained rollout and cyber risk: Release took longer than planned because of safety, security, and alignment work . Access began with Trusted Access Partners, with broader availability contingent on initial results and tiered cyber access based on verification and trust . Astra reached the preparedness framework’s “cyber critical” level, requiring additional safeguards before release . The rationale for releasing cyber capabilities is that increasingly capable models may drive a new wave of attacks while also being needed for collective cyber defense .
- World Labs — Atlas: The team describes Atlas as a newly launched next-generation world model that generates, reconstructs, and simulates worlds. Its core primitive is spatially grounded new-view prediction: users can provide images with camera trajectories and request frames from arbitrary viewpoints. Atlas natively handles text, images, video, camera poses, RGB/depth, and 3D, jointly combining generation and reconstruction rather than forcing outputs through a fixed Gaussian-splat representation.
- Potential productivity and market applications: The team says Atlas can create bullet-time views from three iPhones without studio capture, a green screen, or expensive calibration, and aims to reduce dense 3D reconstruction from roughly 100–300 images to about three—a claimed 50–100× reduction. Proposed applications span film, games, marketing, architecture, construction, and other labor-intensive 3D design workflows.
- Team pedigree and robotics expansion: The team highlights Ben’s background as the creator of NeRF and a researcher who has spent most of his career producing 3D from images. World Labs acquired the company formerly known as Scenix to build real-to-sim and sim-to-real robotics systems; Atlas is intended to reduce the laborious reconstruction and data bottlenecks involved in robotic-policy training. The longer-term thesis is to use learned neural simulators that model how environments respond to actions and potentially turn those simulators into planners. Important caveats are that Atlas currently has only limited “baby dynamics,” while the team identifies training compute—not data or model scale—as the main constraint on further scaling.
- Frontier world models: Runway released Solaris, an “Interface World Model” that generates interactive applications frame by frame in real time—“the image is the application”—without code; its GWM World 2 generates interactive 720p video at 24 fps with 48 kHz audio and responds to user inputs during exploration. World Labs’ Atlas is an omni world model trained from scratch on text, images, video, camera poses, depth maps, and 3D; it generates up to a minute of 1440p video with camera control and reconstructs 3D scenes from one to 100 photos.
- AI for science financing and technical progress: PSI (Physical Superintelligence) raised $58M to advance physics-focused AI. Its team includes Alex Wissner-Gross of Eon Systems, a top-of-class MIT graduate who completed a historical triple major, and Alex Klokus, creator of Gravity Blanket and co-founder of Skywatcher. PSI designed the Fermi Explorer mission after spending a week and 10 billion tokens generating a perihelion-pump maneuver that NASA personnel validated, versus three months and one billion tokens for the initial effort. Arena Physica released Heaviside-1, a second-generation electromagnetism foundation model that the source says can run simulations in milliseconds rather than the hours required by current simulators.
- AI power infrastructure: Fervo signed a 396 MW enhanced-geothermal power purchase agreement with Google for a potential Utah data center beginning in 2028; Google has an option for roughly 600 MW more by June 2030, potentially reaching 1 GW from one customer. Enhanced geothermal uses shale drilling rigs to reach hot rock without requiring a natural hot spring or geyser.
- Physical AI and agricultural automation: Canadian startup Sami Robotics robotically cut and picked an iceberg lettuce head at Reservoir VC’s 40-acre Salinas farm. Reservoir’s stated thesis is backing “rugged AI” startups that solve one high-value agricultural task and expand into adjacent industries; the demonstration targets a particularly difficult automation problem because leafy-green heads vary in size, require precise cutting without bruising, and are already harvested quickly by human crews.
- AI document infrastructure: Reducto says it processes more than one billion pages per month, and its r-1 parsing model combines layout detection, reading order, tables, formatting, grounding, and granular citations in one request; the company claims up to 20% fewer parsing errors, improved high-volume latency, and an all-in cost of $0.01 per page.
- Perplexity CEO Aravind Srinivas described Perplexity Computer as a model-agnostic system that orchestrates compute across models, tools, and chips. Its fully local version runs the model and agent stack on user-owned hardware; released with NVIDIA as “Portable Computer on DGX Spark,” it has attracted interest from finance users seeking air-gapped deployments.
- A hybrid version keeps frontier-model access in the cloud while routing privacy-sensitive material—such as health, financial, and tax records—to a local model. Srinivas said local models still do not match frontier models, making up/down orchestration a practical compromise between intelligence, privacy, cybersecurity, and cost.
- The product points to an emerging distributed-inference thesis: users may contribute local power, memory, and hardware to AI workloads, since Srinivas argues that serving 1 billion always-on agents would require roughly a terawatt of power and cannot rely on data centers alone.
- World Labs’ Atlas is presented as a spatial-intelligence model built on new-view prediction, unifying pixel generation and pixel reconstruction; a16z claims this reduces digital 3D-capture requirements by 50–100x, from 100–300 photos of a room to three.
- The named World Labs co-founders are Fei-Fei Li, Justin Johnson, and Ben Mildenhall, with a16z’s Martin Casado participating in the discussion. An a16z post highlights a founder trajectory from blogging about ray tracing for a 2012 CS project to co-founding World Labs in 2024 and shipping Atlas in 2026.
- The thesis extends to robotics: the discussion identifies data rather than chips as the bottleneck and frames new-view prediction as “AI-complete.”
- Sam Altman describes a founder path combining longstanding AI interest and formal study with an entrepreneurial detour; he says OpenAI was announced at the end of 2015 and began work in January 2016.
- Altman’s technical thesis was that deep learning began working in 2012 and improved with more compute. OpenAI initially explored this through robot-hand Rubik’s Cube and video-game tasks, then made progress on supervised learning and GPT-1 around 2018 and trained GPT-3 in 2020, which he viewed as the first model above a usability threshold.
- He says scaling laws suggested that applying roughly $1 trillion of compute could produce a major capability gain, despite skepticism that such an expenditure was feasible.
- Legora’s revenue has increased by more than 9x year over year despite the company being three years in, signaling sustained growth beyond the typical first-six-month surge.
- Legora reported 36% month-over-month growth in net new ARR; ARR from in-house teams grew more than 15x year over year, average ARR per in-house customer rose more than 2.5x, and in-house teams accounted for nearly half of new customers.
- Monthly active users were spending more than 16 hours in the product, and Legora expected September to add more ARR than the entire first quarter of 2026.
GPT-6 Astra access expanded to all Plus and Business users, following initial availability for Pro, Enterprise, and Business Premium users in Work/Codex and via the API.
AsideAI — AI-agent browser/infrastructure signal. GStack now uses AsideAI as its preferred remote-session browser. Garry Tan says it gives AI agents web and credential access and combines integrations with an agent harness and memory system, calling it his “new favorite AI browser.”
- Venture-backed software’s core assumptions are being questioned, especially which moats will persist and whether software is “over.” The discussion nevertheless frames a barbell future for software around scale and luxury, while emphasizing craft and trust as reasons software is not over.
- The investor-sourcing examples emphasize patience and technical founder pedigree: waiting a year before investing in turbopuffer, and committing to Mesh Optical within 48 hours based on its SpaceX laser engineers despite lacking domain expertise.
GPT-6 Astra is now available to Pro, Enterprise, and Business Premium users in Work/Codex and through the API; rollout to Plus and Business users is next.
- Ollama co-founder and CEO Jeffrey Morgan says the product is used by 9 million developers and 85% of Fortune 500 companies; Ollama Cloud token usage has increased 150× since the start of the year.
- Morgan’s thesis is that open models are less than three months behind frontier closed models, already cover roughly 80% of tasks while being fast and ultra-cheap, and shift the bottleneck toward extreme efficiency.
- The emerging opportunity is software that coordinates multiple inexpensive, fast models to solve complex problems rather than relying on a single larger model, reinforcing an investment theme around orchestration and infrastructure above the model layer.
- Ollama’s reported scale—9 million developers and 85% of Fortune 500 companies—gives co-founder and CEO Jeffrey Morgan a broad view of which AI models people use.
- Morgan identifies a shift toward open models driven by coding agents, falling costs, and capabilities rapidly catching up to frontier labs; Ollama Cloud token usage has increased 150× since the start of the year.
YC affiliation is an important early-stage investor signal: a VC report comparing YC and non-YC companies currently raising found YC startups’ valuations were 79% higher; Paul Graham cautioned that the gap may reflect stronger underlying companies rather than a causal YC effect.
Garry Tan reported that @AsideAI’s AI harness reduced his OpenClaw-with-Slack setup from roughly two hours to under three minutes by combining full integrations, browser integration, and smart access-control defaults.
Open Models Change The Economics of AI
Enterprise open-model demand is moving into production workflows. Jeffrey Morgan says Ollama is seeing a shift toward open models, especially for coding agents and AI assistants; Chinese-origin models currently dominate Ollama Cloud usage and are accessed by businesses worldwide, with the US and Germany prominent sources. AT&T reportedly shifted 40% of its token consumption to open models while evaluating Chinese models. The interview reports roughly 150× growth in cloud token usage since the start of the year, first driven by coding agents and then by OpenClaw and Hermes, which extended long-running automation to non-developers.
The strongest model-economics theme is cheap, composable open models rather than a single “god model.” Morgan forecasts that 80–90% of enterprise tokens could run through open models while only 10–20% of spend goes to them, with frontier models reserved for the hardest tasks and routers combining open and closed models. Rapid model-release cycles make custom training harder, although tooling is improving. DeepSeek Flash is presented as a fast, ultra-low-cost model good enough for roughly 80% of tasks and the highest-growth model family on Ollama Cloud; chaining such models through a coordination layer is described as an opportunity for new workflow startups.
The investable infrastructure gap is above the model-token layer. Supporting each new open model requires compatible inference engines, harnesses, benchmarks, cloud capacity, hardware integrations, and a common runtime. The emerging layers include company knowledge/context, multi-agent coordination, execution sandboxes, stateful storage, credentials, safety, and curation across fragmented providers. Startups also face high GPU-access friction and volatile supply and demand for B200/B300-class hardware.
Ollama provides a strong founder-pedigree and product-traction case. Its cofounders previously built Docker Desktop at Docker, were second-time founders, joined YC in 2021, and pivoted from Kubernetes/security work to locally hosted LLMs. They say they partnered with Benchmark in 2022 around a Series A pitch focused on developer experience and security. The interview characterizes Ollama as having 9 million developers, 178,000 GitHub stars, and usage across 85% of the Fortune 500.
Security, safety, and provenance remain adoption risks. Customer discussions identify security and safety as major blockers to open-model adoption, while open models lack some safety tooling provided by closed-model vendors. Customers also weigh where a model runs, its origin, and data provenance, especially for mission-critical applications.