We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
OpenAI's math release: big claims, still under review
OpenAI published mathematical results from an internal frontier model it has not released. It says it consulted the Institute for Advanced Study's Advisory Group on Mathematics and AI on how to release them . Reported figures:
- Size: 722 manuscripts, grouped into 372 families of related results.
- Source: an evaluation of about 4,000 research problems.
- Compute: roughly three hours of ChatGPT Pro thinking per result, on average .
- What's included: papers, proof artifacts and selected reasoning summaries .
Sam Altman called it "a new era of discovery" .
Reaction is split. Mathematician Levent Alpöge called it "the most significant moment in mathematical history" . One analysis estimates about 20% of the results are disproofs or counterexamples . Will Depue expects some results not to survive scrutiny. François Chollet asks whether progress in verifiable domains like math and code carries over to domains that still depend on human data . Individual highlights, such as integer multiplication faster than n log n, have not been independently verified .
Why it matters for investors: treat this as a claim about scientific capability to watch over the coming weeks, not a settled result. The open question for AI-for-science companies is whether these gains extend beyond domains where answers can be checked automatically.
Mistral Large 4: a strong preview with disputed benchmarks
Mistral previewed Large 4:
- Size: 1T total parameters, 49B active, natively multimodal .
- Availability: API now, open weights promised for the end of October. Mistral says it is "forged in Europe end-to-end" and is working privately with cybersecurity partners .
- Training: about 3,800 Grace Blackwells in Europe. The RL run is still going and shows no sign of saturation, and a larger model is already training .
- Pricing: $1.36 per million input tokens and $4.18 per million output tokens .
Independent results are mixed:
- Artificial Analysis: an Intelligence Index score of 38, the top score from outside the US and China. But cost per task is more than 4x that of open models with similar scores .
- Cyber benchmarks: Cline attributes much of the cyber lead to the model refusing fewer tasks .
- Comparison with GLM: critics note it trails GLM-5.3 on Artificial Analysis's index .
- Open-weight label: Clem Delangue pointed out that it can't be called the best open-weight model until the weights actually ship .
For investors, the takeaway is that Europe now has a credible sovereign model built on its own compute. Its cost competitiveness against Chinese open models is not yet established.
AI is producing far more security findings
a16z data shows reported critical vulnerabilities at 21 major software companies, including Apple, AWS, Microsoft and Google. They never exceeded 100 a month in four years, but have topped 600 a month since spring . About 87% of exploited bugs are now attacked on or before the day they become public, up from 23% in 2020 .
Armadin. Kevin Mandia, who built Mandiant, says Armadin has found more than 90 zero-day vulnerabilities at customer sites since January, all in production systems. It works as an outside "black box" attacker with no source code, using models post-trained with real exploit developers . By his own account, Armadin has no sales force or go-to-market strategy yet, and needs funding to build them .
OpenAI's internal effort. Greg Brockman said OpenAI moved 25% of its production engineers onto security, using models to find holes. That work eventually "saturated" at the critical issues Astra could find. OpenAI is now building an automated "defense factory" to rerun the process with each new model .
The investable idea is continuous, model-driven offense and defense that repeats with every new capability release.
Agents meeting businesses: a protocol, and a privacy problem
Meta and Sierra announced the Personal Agent Protocol, an open standard for how personal agents interact with businesses. Partners include Genesys, Instinct, Shopify, Stripe and Walmart . Sriram Krishnan compared it to OAuth. He asked how Amazon, airlines and other aggregators will respond if they can no longer own the final user experience . This picks up the question from the last brief about where network effects end up in agents.
In the same period, Time reported that Meta's Muse keeps hourly-updated dossiers on users and the people they mention. Lessons learned are shared across Muse agent instances, even though Meta describes each user's virtual machine as isolated .
How engineering work is changing
The Pragmatic Engineer reports that the most productive engineers now run 5–10 agent sessions in parallel and have stopped writing code by hand . Agent-authored pull requests on GitHub outnumbered human-authored ones in August . At Linear, agents have created more issues than humans since July .
Costs are a counterweight. Uber's token usage rose while its costs stayed flat, thanks to open models and smart routing between models . On capacity, Google's Amin Vahdat says Google aims to double effective token-serving capacity every six months, with as much or more of that coming from software as from hardware .
Support-agent benchmark. Gorgias, a vendor that ranks fifth in its own test, benchmarked 13 AI support vendors on 212 live stores. The top five resolve 64–75% of conversations with no human involved, and the median resolves 48% .
Deals
- Valon raised a $150M Series D at a $2.3B valuation. It signed more than $200M in deals within six months of starting software sales. It first ran its own mortgage servicer until it reached 3x industry efficiency. One of the largest US servicers is moving 4M loans onto its platform, nearly 10% of the market .
- General Medicine: a16z led a $120M Series B for this "general store for healthcare." It combines pricing, clinical guidance and a clinician marketplace, and is also building for enterprise customers. It reunites the PillPack team .
- Multiply Labs (YC S16) raised a $75M Series B for robots that manufacture gene therapies and mRNA drugs. Its robots run inside AstraZeneca and Legend Biotech facilities .
- Type One Energy raised $200M for fusion. Its CEO says this covers about half of a 400-MW plant targeted for 2034, built through a project-specific supplier network .
- Outer Space launched with an $8M pre-seed led by Upfront. It builds outdoor structures that generate and store solar power for homes .
Investor debate and signals
Seed checks as options. Leo Polovets responded to Menlo's view that seed checks are options to be sized up later: "No seed investment should be an option bet." His fund does little follow-on investing by design .
Software moats. Naval argues models are software's last moat, as AI decompiles and recodes products. He expects more software to retreat to the server . Amjad Masad predicts all software will soon be "de facto open-source" .
Robotics. Generalist says its Gen 1.5 robot model improvised with unfamiliar tools, using a banana as a brush and a dustpan with both hands. The company says it was not explicitly trained to do this .
- The interviewer cited more than $200B in Google capex expected “this year,” mostly for data centers, and called the buildout historic; Vahdat agreed its scale and pace were unprecedented. Vahdat identified power as the most fundamental long-term constraint: Google prefers utility partnerships planned years ahead, says it covers the grid-infrastructure costs it triggers, and may use local generation or batteries to bridge supply gaps.
- Google’s target is to roughly double effective serving capacity, measured by token generation, every six months—not necessarily to double FLOPs. Vahdat said software and model improvements may contribute as much or more than hardware, and most observed intelligence-per-watt gains had come from model/software improvements.
- Long-horizon agents can issue follow-up requests in milliseconds rather than waiting for human-paced interactions; their CPU-based reasoning and context gathering also increase demand for CPUs, networking, and storage alongside accelerators.
- Google evaluates workload-specific “goodput”—useful completed work—rather than theoretical FLOPs; at 100,000 accelerators, Vahdat said failures could occur multiple times daily or hourly, with detection and recovery affecting delivered performance.
- AI data centers are increasingly co-designed around their hardware: accelerator racks can draw hundreds of kilowatts versus roughly 10–40 kW for storage racks, affecting power, cooling, and network design. Google also described separate 8i inference and 8t training chips as inference demand grew; both can run the other workload, and specialization must justify its fixed cost with durable demand.
- Google says hardware engineers now use AI as much as software engineers, with productivity gains and shorter design-to-tapeout and bring-up times; AI also helps assemble information for data-center planning but is not replacing human judgment.
- Google is pursuing orbital data centers as a “moonshot”; Vahdat cited 98–100% sunlight coverage in sun-synchronous orbit versus about 28–35% on land, while noting that cooling, reliability, and repairs are harder in space.
- Kevin, who says he built Mandiant and spent 30 years in security, says he joined Armadin after meeting founders David Slater, Travis Lam, and Evan Peña, whose team and proposed product convinced him the company addressed an urgent need created by AI’s impact on security. He argues that open models are already capable enough for structured vulnerability exploitation, and predicts that cheaper, more anonymous access to compute will expand criminal AI-enabled attacks, improve less-skilled attackers’ effectiveness, and make attribution harder.
- Armadin Red is positioned as continuous AI-enabled offensive security: agent swarms map network assets and services, monitor changes, and retest them; the company says it aims to verify exploitable risk rather than return only a list of known vulnerabilities.
- The interviewee reports that Armadin found more than 90 zero-day vulnerabilities since January 2026 in customer production environments, including at Fortune 500 companies, using outside-in, black-box testing without source-code review. He says most of those findings were discovered by humans while AI performed over 90% of routine work, but that the technology itself found their latest zero-days.
- In an Armadin test of 20 human-executed attack chains, the interviewee says no tested model completed more than eight; open-weight and leading closed models reached the same result, with closed models faster but open models catching up when given more time. He describes cost as the remaining differentiator.
- Armadin Blue is intended to turn exploit findings into compensating controls through EDR and firewalls; the interviewee described its first controls as forthcoming and initially limited, and said Armadin was working with CrowdStrike and Palo Alto Networks. Commercialization remains a stated gap: he said the company had no sales force, go-to-market strategy, or international operation yet, and that funding and scalable operating leadership were needed to build them.
- Generalist CEO Pete Florence says Gen 1.5 shows one- and few-shot learning across varied tasks and zero-shot physical generalization, capabilities he says were not explicitly trained for. In demonstrations, a robot trained to sweep with a brush used a banana as one; when given a dustpan, it used its other hand to sweep a cube into it, then tilted the dustpan to dump the cube into a bowl.
- Florence describes Gen Zero as establishing predictable robotics scaling laws as compute, data, and model size increase. He says Gen One was shown achieving 99%+ success rates on several tasks; his definition of mastery also includes speed, reliability, and improvisational responses to unexpected situations.
- Florence says Gen One had zero robot data in pretraining and achieved high success on varied tasks with one hour of data. Generalist says it is scaling data across more than 9,000 physically different hand configurations, including two- and five-finger hands and specialized tools; generalizing to new hands remains not fully solved.
- Generalist works with partners to deploy models in real environments and use real-world evaluations to identify capability gaps and guide research priorities. Florence says partners also contribute knowledge of relevant machines, processes, and workflows, while this work requires substantial relationship-building and deployment effort.
- Altara targets physical-science AI in semiconductors, batteries, and advanced materials, positioning its product as a bridge from fragmented experimental and production records, SEM images, wafer maps, simulations, and spreadsheets to experiment and production decisions. Altara described a semiconductor-value-chain customer analyzing historical R&D and production data, simulating a new experiment on Tuesday, and running it at a plant on Wednesday.
- Altara’s co-founder sees AI for science moving beyond helping researchers understand science toward helping them do it, with robotics, simulation, and agentic tools expanding AI’s role; she identifies scarce high-quality data, unresolved closed-loop experimentation and scientific evaluation, and real-world deployment as open challenges, while arguing small teams can accelerate R&D and commercialization.
- Altara’s founding engineer used semiconductor etching to illustrate the verification challenge: the process is costly, irreversible, and multistage, errors can propagate, and wafers are scarce; he described physics-based digital twins and simulation as potential verifiers. In his framework, falling offline-RL loss and stabilizing rewards do not establish physical-world performance; stronger checks compare with real-world data and introduce an overetch at one stage to test whether the policy changes downstream recipe and process conditions.
- AI use is shifting from assistance to delegated work: engineers interviewed describe running roughly 5–10 agent sessions in parallel, and the author says nearly all highly productive engineers he met who had hand-written code a year earlier had stopped doing so. Linear reports agent-created issues outnumbering human-created ones since July; GitHub data is described as showing rapid growth in agent-authored PRs, which the author says exceeded human-authored PRs in August. Factory AI reports that 83% of its users use agent skills, up 2.5x from February.
- Internal agent infrastructure is becoming common: the author says most mid-sized-and-larger companies have built their own harnesses, citing examples including Stripe, Uber, Shopify, Google, Meta, and Amazon. Development increasingly starts in Slack and triggers cloud agents, while agentic software factories connect agents to systems such as CI/CD, code review, and deployment; the author expects cloud coding agents and harnesses to expand, including offerings from vendors such as Anthropic and OpenAI and new startup products.
- AI software economics have two opposing signals: companies are adopting open models and smart routing to reduce token costs by 50% or more; at Uber, token use rose while costs stayed flat from May. Meanwhile, the author reports GPU and memory shortages alongside tighter CPU supply, with spot access difficult and some CPU reservations needed months ahead.
- More generated code is straining quality controls: an engineer at a mid-sized startup describes human reviews becoming “LGTM” rubber stamps under the volume, and Linear data shows AI-only PR reviews increasing. The author identifies agentic evaluation in CI/CD and agentic observability across deployed systems and incident management as emerging infrastructure needs. As a reliability caveat, Pi creator Mario Zechner says software feels increasingly brittle, while acknowledging that some of the problems predate agents.
- Leaner project staffing is emerging, but not the disappearance of engineering teams: Anthropic’s Head of Engineering says projects often have no more than two engineers because each is already managing agents, and the author says startup teams are shrinking; she also says teams continue to own software and on-call responsibilities and remain similar in shape to two-pizza teams. One Series D startup says it screens for AI-positive engineers. The author found no company where non-engineers were shipping production code; PMs and designers were prototyping or sending changes to engineers, who retained release responsibility.
- Altman described Dots as a personal AI helper that handles unwanted email and to-do work; in one example it evaluated ideas, found material on his computer, and proactively contacted three teammates. He said privacy must be “extremely” rigorous for the product to work.
- Altman said the planned Astra 6.1 release was held back because the ready model leaned too far toward accident risk; OpenAI is pursuing a model it can still distribute broadly while making stronger alignment claims.
- Altman supports open-source models despite warning of a coming wave of cybersecurity problems and saying society may have to accept fairly severe incidents in exchange for the liberty they provide. He also favors a pro-innovation approach with guardrails, pre-release standards and liability; he said safety assessment now needs to cover training, when models might escape a sandbox or hack systems.
- Reflection launched Beam, a text-only 501B-total/23B-active MoE trained from scratch for coding, agentic and scientific work, with full Apache 2.0 weights promised later that month. Team posts cite 23.8T pretraining tokens and RL/OPD training on 10K GB300s with over 100M rollouts across about 1M tasks; Reflection claims 80.9 on SWE-bench Verified and 3–4× the inference efficiency of GLM 5.2. The newsletter says newer leading Chinese models are generally ahead, with observers placing Beam around GLM 5.2 and below DeepSeek V4 Flash on some benchmarks; it nevertheless identifies demand for US-trained alternatives and calls the launch Reflection’s arrival as a functional neolab. Axios reported Reflection pays $150M/month for Colossus compute and has a $1B Nebius deal.
- Hugging Face’s multi-harness RL proxy can turn 10 unmodified agent harnesses into RL environments. The same weights scored 62% with Mini-SWE-Agent versus 33% with Claude Code; training across four harnesses raised first-attempt solves from 42% to 54%, while a tool-call bonus cut calls by 31%. Results are limited to one task family and one seed.
- Two model releases illustrate specialization: Reka’s Rho-1 is a 19B model for text, images, video and robot actions, trained from scratch on 320 H100s in about three months; Command Code’s Agr (31B) and Agr-flash (360M) return typed values and per-option probabilities for tool calls and routing rather than generating text.
- SemiAnalysis found Claude subscriptions delivered 5×+ the API-equivalent value of OpenAI plans, but adjusting for task cost narrowed the advantage to 1.3–2.9×; its methodology uses model- and token-specific credit costs, not list API prices. Users also reported paused new $200 sign-ups and effectively halved OpenAI usage limits; the newsletter reports default speeds for GPT-6 Astra and GPT-6.1 Sol subsequently rose about 50% across subscription surfaces and partners.
- AMD reportedly bought World Labs for $8.2B.
- Robotics models are moving from action-by-action imitation learning, which transfers poorly across actions and embodiments, toward visual-language-action models; emerging world-action models predict a future world state and actions to reach it, with broader generalization still a thesis. Robotics training also lacks an internet-scale corpus: companies use teleoperation, egocentric recordings and simulation, while learning from a robot’s own real-time interactions could provide higher-quality, embodiment-specific data.
- Deployment remains a potential moat: real-world robotics is difficult and operationally intensive, and companies with customer relationships and deployment experience may be able to adopt more general models later. The investor sees room for deployment-focused companies rather than only heavily capitalized foundation-model builders.
- Humanoid cost declines and hardware commoditization are not assured: actuators reportedly account for 40–60% of build cost, and actuator and rare-earth-magnet supply chains depend heavily on China amid tariffs and restrictions.
- Robotics data-collection businesses and robot-specific inference/physical-world chips are already emerging; data-quality tooling, observability and standardized evaluations are prospective infrastructure opportunities. Current evaluations are hard to compare across companies and matter for customer trust and safety. The generalist-model quadrant is described as the most capitalized; founder differentiation could come from architectural gains, proprietary data collection or greater computational efficiency.
- The guest’s thesis is that Meta, OpenAI, and Anthropic may evolve beyond LLMs into biotech, with non-invasive brain read/write as a shared goal; he sees precise neural modulation, with AI learning to titrate individual brain states, as a major frontier.
- An unnamed MIT team has developed an eye mask that directly measures REM and uses behind-ear stimulation to induce eye movements; the guest said it helped him fall asleep faster and amplified REM, and estimated release in 7–12 months.
- An unnamed neuroscientist and biotech founder is developing a cuff to directly measure autonomic nervous-system activity and distinguish distress from positive arousal; the guest said he had tried it and that it had been brought to police and first responders, with broader availability envisioned eventually.
- Technical caveat: the guest said current TMS and ultrasound stimulation lack good spatial control, while memory and thought remain difficult to model and neurons participate in multiple behaviors—complicating precise neural targeting.
- Altman says OpenAI paused a model-training run when capabilities were advancing faster than alignment and monitorability, and wants to move to each new capability level only when safeguards keep pace. OpenAI is discussing safety with other labs and governments; he says verifiable shared U.S.–China safety thresholds would be a major step.
- For a possible IPO, Altman says OpenAI is not in a rush and its nonprofit will continue to govern it; he says the company would prioritize its mission over shareholder returns or revenue, even if that makes its stock volatile.
- Altman expects AI to advance treatment of some diseases within several years and substantially change health care within a decade. He pointed to college students building a self-driving golf cart with AI-assisted engineering as an example of more ambitious projects becoming feasible, and forecast a major boom in entrepreneurship.
Garry Tan reports using Opus 5.5 in fast mode to implement Doom in Paul Graham’s Bel Lisp variant in about 20 minutes; the result comprised 32k C++ transpiled to under 2k lines of Bel, plus a full Bel interpreter in JavaScript. In a follow-up, Tan said maybe three more prompts brought it to 35 fps at 640×480, using eight CPU cores per frame. He shared a playable demo and source code.
Garry Tan says his GBrain bug-fixing workflow is chaotic because he opens a new thread per issue, but calls Capy’s cross-thread coordination and “virtual standup” conflict avoidance on parallel threads “SOTA.” He says GStack’s autoplan and eli5 help him move faster.
Scott Kupor recommended U.S. Tech Force and EarlyCareers.com as starting points for people seeking public-sector roles, amplifying Sriram Krishnan’s call to connect more tech talent with public-sector leaders.
Garry Tan says agents can use markdown skills and cron jobs to perform “almost every useful type of knowledge work.” He amplifies Meta chief AI officer Alexandr Wang’s account that, in some internal Meta cases, agent swarms with the right agent loop and evaluation metric can outperform a team of 100 engineers.
- General-purpose agents are not inherently winner-take-all: different agents can interoperate through messaging, APIs, and other channels, while portable text-based memory and context can ease switching. Better models, UX, integrations, and distribution may still support large companies, but Chen distinguishes those advantages from network effects.
- Potential agent-specific network effects include turning generated artifacts—such as shareable plans, microsites, or group workflows—into acquisition loops, and building shared identity, trust, reputation, private context, and intent that improve multi-party matching. Such viral loops depend on access to contacts and communication channels, plus user trust.
- The strategic fault line is where those network effects reside: open protocols and specialist networks could leave agents interchangeable and allow vertical agents to thrive, while horizontal agents that retain identity, relationships, context, and distribution could make switching mean leaving a network behind. Chen expects a contest between horizontal agents seeking to internalize network functions and incumbents or startups seeking to make their networks agent-accessible.
Paul Graham’s view is that money-motivated competitors are either ineffective or, if capable, eventually “get it” and stop working—a general observation about competitor incentives, not a company-specific development.
Scott Kupor says the “SI Force” will work to reconcile the administration’s goals of extending the U.S. lead and doing so responsibly .
Sam Altman said “We are entering a new era of discovery now” and linked to OpenAI’s page titled “Sharing AI progress in mathematics,” signaling a mathematics-focused AI progress update without further details in the post.
YC S16 startup Multiply Labs raised a $75M Series B to build robots that manufacture complex drugs, including gene therapies and mRNA treatments. The company targets a manufacturing bottleneck as AI speeds drug discovery: production is still mostly manual, slow, and vulnerable to contamination, and some gene-therapy doses can cost more than $1 million to produce. Its robots run production inside facilities at AstraZeneca and Legend Biotech.
Of the software bugs hackers actually exploit, about 87% are now attacked on or before the day they become public, up from 23% in 2020—evidence that the patching window is shrinking sharply.
Kevin Mandia describes Armadin’s approach as using agent swarms to map networks and poll for changes like a heartbeat; the company’s stated goal is to determine whether a known vulnerability is exploitable before attackers do.
The state of the tech industry in 2026
I recently delivered the keynote at the LDX3 engineering leadership conference (opens in new tab) in New York, attended by 2,000+ engineering leaders, CTOs, Director+ and Staff+ engineering folks, in which I attempted to create a snapshot of where the tech industry is at this exact point.

In the middle of my keynote at LDX3
It came after I had the chance to visit the AI labs OpenAI and Anthropic, meet innovative startups like Ramp, Uber and others, get new, unpublished data from GitHub, Factory AI and Linear. The focus of the talk is trend inside of AI labs, VC-funded startups and Big Tech. Thanks to the kind access to these teams!
Before the conference, I spent weeks identifying and recording trends that didn’t exist a year ago – or were much more nascent back then – and also things that are currently broken, and others that remain unchanged. Today, we cover:
Speed of change. In the tech industry, change has never been this large-scale, this fast. Industry legend Martin Fowler confirms.
What’s changed: practically nobody writes code by hand anymore, working with 5-10 agents more common, the fading of the IDE, and 13 more changes from the last year.
What’s broken: assumptions about code output broke, code reviews became theatrical, quality and reliability is down, and more.
What’s still the same: teams are still important, as is planning; non-engineers are still not shipping code, and more.
What comes next? Cloud coding agents and harnesses, engineers to stop reading the code, companies building a new type of AI infra, and more are already starting, and should accelerate.
You can watch the full talk, which is 29 minutes long:
The bottom of this article could be cut off in some email clients. Read the full article uninterrupted, online. (opens in new tab)
1. Speed of change
People in tech are more than used to change; seeing the internet go from zero to everywhere, the arrival of mobile communications and smartphones exploding in popularity, cloud computing rising from a niche concern, and the development of programming languages like Go and Rust, and frameworks like React (web) and Jetpack Compose (Android).
Despite this, the scale and pace of the impact AI is having at present is still without precedent. That’s how industry veteran Martin Fowler described it at The Pragmatic Summit (opens in new tab):
“Nothing has hit with the magnitude of AI. This is a whole size difference from anything that we’ve faced before.
On a smaller scale, we were very much involved in the growth of object-oriented languages. [object-oriented languages] scared a lot of people, but it didn’t scare us so much because we were part of it.
The internet had a huge impact upon us all. And, of course, we were spreading the challenge of agile software development. [Agile] had a very big impact on a lot of organizations because you could tell by how hard they resisted it.
But [for all those other changes] we kept talking about how important they were and how valuable they were and trying to persuade people of the importance of them. It may sound surprising, but even for the internet, there were people who wouldn’t think that was important!
For AI there’s no argument about how important it is. You cannot put blinkers on to deny the importance of this thing.”

Martin Fowler (left) at The Pragmatic Summit. See the full video (opens in new tab)
My own take is similar; we are without doubt at the beginning of a massive technological transformation, and after it, software will be built differently from how we’ve been doing for decades, although some parts of development will remain the same. At the very least, the tooling and best practices will look different, and things are changing faster than ever.
2. What’s changed?
#1 Nobody writes code by hand anymore
As the calendar enters Q4 of this year, we see this has solidified into a mega trend since the end of 2025 year when models greatly improved at coding. Related to this, in the very first article of this year we asked the question: when AI writes almost all code, what happens to software engineering? (opens in new tab)

Today, there are plenty of signs that most engineers have stopped writing code by hand, and it was also a sign of the times recently when the Ruby on Rails creator sparked a “death of coding by hand” debate. Personally, I’m excited by the new opportunity that AI creates, and it’s exciting to experience this revolution firsthand. As mentioned in that deepdive (opens in new tab), I’m certain that my way of coding will change drastically in 2026, and there’ll be plenty of knock-on effects.
For more on this topic, check out this recent edition of The Pulse (opens in new tab).
#2 Working with several parallel agents is increasingly common
I like talking to engineers at AI labs because their working practices are often a few months ahead of the rest of the industry. Claude Code creator, Boris Cherny, revealed the new way he gets things done on the podcast: (opens in new tab)
“I have 5 terminal tabs. Each one of them has a checkout of their repository. I’ll round robin and start Claude Code in each one. I also run 5-10 Claudes on Claude Web, in parallel with my local Claudes.”
A few months after these comments, on last week’s pod (opens in new tab) with Cockroach Labs cofounder, Peter Mattis, – who’s an extremely productive developer who wrote circa 100K lines of code/year, pre-AI – said he’s doing something similar:
“Oftentimes, I’m doing things in parallel. I find my cognitive overhead is about 5-10 agent sessions concurrently. But sometimes those sessions will have many subagents doing things.”
I also asked a former coworker at Uber, who’s also an extremely productive software engineer, how he works these days. Dima Zaytsev (now a software engineer at Linear), told me:
“It used to be a simple world: one mouse, one keyboard, one screen. You physically can’t work on more than one thing at once. Now, you no longer have that limitation.
I end up having 5-10 (literally) worktrees locally that I rotate between. Prompt one agent, and while it is working, move on to another to test output or review code.”
Nearly all of the most productive software engineers I’ve met – who were hand-writing code a year ago – no longer write the code by hand and also run several parallel agents. What a change!
#3 Rapid rise of agent-generated PRs
Fresh data that GitHub shared with me:

Agent-authored PRs: ninefold increase in eight months’ time. Source: GitHub (opens in new tab)
I’ve given GitHub grief for its ongoing reliability issues (opens in new tab), but seeing how rapidly agent-generated PRs are growing makes me a lot more empathetic for the load which the platform is dealing with.
Food for thought: there were more agent-authored PRs in August 2026 on GitHub than human-authored ones! It’s reasonable to speculate that the number of agent-generated PRs will be permanently exceeding human-generated ones on the platform, going forward.
To give a sense of the number of AI-generated PRs: as of today (October 2025), there is likely to be 3x as many fully AI-generated PRs, per month (probably 75M+), than human generated ones at the end of 2023 (25M in December 2023). And the pace of AI-generated PRs does not seem to be slowing down, at least not now.
#4 Agent-generated issues also rising rapidly
Agents are not just generating PRs, they’re also opening tickets. Some more exclusive data, this time from my friends at Linear, which added first-class MCP support (opens in new tab) for agentic creation of issues:

More agent-created issues than human-created ones since July. Source: Linear (opens in new tab)
#5 Agent skills usage is mainstream
Both engineers and non-engineers use agent skills, and lots of companies are building their own agent skills repositories for all staff to create, share, and evaluate. An interesting data point on skills usage trending heavily up comes from autonomous software factory vendor Factory AI (opens in new tab), which shared with me:

83% of Factory AI users use skills, up 2.5x from February. Source: Factory AI
#6 The fading IDE
I have great respect for software engineer Steve Yegge, who’s good at identifying emerging trends and learning about them. In a podcast episode (opens in new tab), he said the IDE is effectively over:
Gergely: “Speaking about the job as developers, you’ve said something that can be triggering for a lot of people. You’ve said, on the AI Engineer Summit, that if you’re still using an IDE now, you’re a bad engineer.”
Steve: “Yeah, well, you got to be a little provocative. Let me put it this way. Okay, I’m not going to say you’re a bad engineer because I know some very, very good engineers, better than I am, who are still at level one or two in my chart [of AI usage]? But I feel profoundly sorry for them.
I feel pity for them like I’ve never felt in my life. For these grown people who are good engineers or used to be. And they’re like, ‘yeah, you know, I use Cursor and I ask it questions sometimes. And I’m really impressed with the answers. And then I review its code really carefully.’
And then I check it in and I’m like: ‘dude, you’re going to get fired. And you’re one of the best engineers I know!!’ “
Steve described his levels of AI usage like this:
No AI usage
AI agent in the IDE, with strict permissions
AI agent in the IDE, with “YOLO” mode (permissions off)
AI agent in the IDE, no longer looking at the code, but interacting with the agent
CLI-first: abandoned the IDE
Running several agents in parallel
Running 10+ agents
Built a custom agent orchestrator to run 30+ (or 100+) agents
Steve believes that “AI-pilled” engineers will abandon the IDE, and I’m seeing data from the market which backs that prediction:
Antigravity 1.0 was the last IDE released, based on a VS Code fork. It was released in November 2025. After that, no major IDE that is a VS Code Fork has launched. In May 2026, when Antigravity 2.0 launched, the product moved away (opens in new tab) from the IDE concept.
Codex launched as a non-IDE in Feb 2026, and they’re glad they did so, despite being torn about (opens in new tab) moving away from the IDE concept, at the end of 2025.
Cursor moved away from the IDE. In April 2026, Cursor relaunched (opens in new tab) and got rid of the IDE interface. It now looks pretty similar to the Codex and Claude desktop apps. When I talked to their team over the summer, they told me they maintain the “legacy” VS Code fork version for existing enterprises, but the future is agentic for them.
JetBrains seems in a hurry to pivot from IDEs. JetBrains has built some of the most beloved IDEs for developers, but the company is now also pivoting to JetBrains Air (opens in new tab), and agentic development environments.

How every coding tool looks these days – gone is the IDE? Source: Cursor (opens in new tab)

JetBrains also joining the trend with JetBrains Air (opens in new tab)
Industry legend Kent Beck told me something interesting when we discussed why IDEs are slowly vanishing:
“It’s not that we don’t need more perspective and context for our human-based decisions; it’s that the context has changed.”
I wonder if what is really happening is the IDE evolving into something different to work better with coding agents. Steve Yegge also talked about how the IDE needs to evolve into a conversation and monitoring interface, but not an editor, given that we no longer write the code.
I’d also add that IDEs need to evolve into a validation and verification interface: how do we know that the agents’ output works as expected, how can we verify what tests it passes, how it looks, and whether you can try out whatever the agent builds. Tools like this are surely coming and will be widely adopted.
#7 Everyone is building their own harness/agent platform
In August, I posed a question on social media about whether it’s possible to be a serious tech business without having built an AI harness:

A provocative question (opens in new tab)
I asked it because most mid-sized-and-above companies have built their own agent harnesses, at this point! A few examples:
Ramp building Inspect (opens in new tab) – a case study we covered in depth (opens in new tab)
Stripe (Minions), Uber (Minion), Block (Goose), Shopify (River)
Google (Agent Smith), Meta (Devmate), Amazon (Kiro + Crew), Dropbox (Nova), Spotify (Honk)
DoorDash (Flux), Grab (LLM-Kit), WorkOS (Horizon), Hubspot (Crucible)
Monzo (Agent Chip), Sierra (Pinecone), Harvey (Spectre), Browserbase (bb)
… and many, many more!
#8 Dev work often starts in Slack

LDX3 keynote: much dev work starts in Slack for startups
What I’ve learned by visiting startups is that more and more development kicks off inside Slack. At OpenAI, it’s “@Codex, implement this”; within Anthropic, it’s “@Claude build this”, while inside Linear, it’s “@Linear work on this”. More startups have their own Slack integration of their coding agent – often a custom one.
Coding agents work really well when deeply integrated into a company’s stack, when the Slack agent kicks off a coding agent that runs in the cloud.
#9 Agentic software factories being built everywhere
We have covered in-depth how OpenAI is building an agentic software factory (opens in new tab), and a lot of feedback from readers about that article has been an “we are too!” type response. Many readers say that their company is building a similar “agentic software factory.”

From the deepdive Inside OpenAI’s agentic software factory (opens in new tab)
The “agentic software factory” adds AI functionality to a bunch of existing systems such as the CI/CD system; for instance, in being able to interact with CI/CD via Slack, or creating brand new systems with agents in them like agentic code review tools, OpenAI’s agentic deploy system, or the Perf Factory.
#10 Migrations no longer take years
We’re seeing many migrations that would have taken years to complete, now taking mere weeks or months:
Anthropic: migrating the package manager Bun from Zig to Rust in 11 days (opens in new tab) (versus an estimated 1.5 engineering years)
OpenAI: the API layer is being migrated from Python to Rust. The migration is taking around 5 months (opens in new tab), versus several years that it would have taken, without AI
Airbnb: migrating 3,500 end-to-end test files from Enzyme to React Testing Library: took 6 weeks (opens in new tab) (instead of years)
Asana: migrating 4,000 end-to-end test files from Enzyme to React Testing Library: 2 weeks (opens in new tab), instead of an estimated spread out over five years
Uber: migrating 600,000 JUnit 4 tests across 15 million lines of code (!!) to JUnit 5: took 4 months (opens in new tab) (instead of estimated several years)
#11 AI costs a major engineering concern
In May, we covered the trend of companies wanting to cut back on AI spend within engineering departments (opens in new tab), and last month, I reported on how tech companies are moving to open AI models to save 50% or more in token costs (opens in new tab). Uber is a good example, where, despite token usage trending upwards, costs have stayed flat since May thanks to them running open models and doing smart model routing:

Uber’s per-token costs are down 50%+. Source: The Pulse (opens in new tab)
Three weeks after my article, Bloomberg confirmed (opens in new tab) the exact same trend. As I covered in my original deepdive, lower-cost models and smart routing account for the majority of cost savings at most companies. If you’re only doing two things, consider those techniques. (opens in new tab)

Most efficient way to cut back on token costs. Image source: Databricks (opens in new tab)
#12 Projects get done with only one or two engineers
Katelyn Lesse, Head of Engineering for Claude Platform, told me how her team works inside Anthropic. From our deepdive (opens in new tab):
“On an individual project, you often cannot have more than two people working on it. This is because each engineer is already running several agents. And so as an engineer, you’re already fighting against your agents, which are stepping on each other’s toes on implementation. And in this setup, you just cannot have that many humans, who also come with all their agents!”
These days, I see the same thing happening at companies both small and large.
Other changes
13. Engineering specializations are disappearing. Specifically, there seems to be less demand for iOS, Android, and frontend specializations, and more for “generic” software engineers. We covered this trend with data in the State of the software engineering job market in 2026 (opens in new tab), observing how mobile and frontend demand is dropping (opens in new tab) (while AI & FDE demand surges)
14. Teams are getting smaller. Projects are being done by fewer engineers. Inside startups, team sizes seem to be shrinking.
15. Reduced junior recruitment. We cover how it’s harder for graduates & interns to get hired (opens in new tab) in the State of the software engineering job market in 2026 (opens in new tab) deepdive.
16. Agentic infra is becoming its own discipline. Most mid-sized-and-above companies are building out agentic platforms, and infra teams are increasingly morphing into agentic infra teams.
3. What’s broken?
Assumptions about code quantity and frequency
It’s been a commonly held belief that code quantity grows roughly linearly. But with agents, both the lines of code and number of commits are growing exponentially, as shown in data from GitHub:

(Human) code reviews are dead
A software engineer at a mid-sized startup told me something that’s pretty much an open secret across the industry, including at places with code reviews in place.
“Everyone is playing the theater of doing reviews, but with the volume of changes that get thrown your way, I observed that people just find a path of least resistance: give up on reviews and just stamp everything with LGTM.
We are gradually phasing out this process ourselves with some shortcuts that you are allowed to take, and that has been a major productivity boost.”
At the LDX3 conference, there was a strong reaction among the audience when they heard the bolded sentence out above. Today at most companies, there is indeed a “theater of code reviews”. The fact is that no engineer can keep up with 5-10 times more code to review, so most don’t thoroughly review the code.
Agent-only code reviews trending up
Fresh data from Linear shows that PRs which are reviewed solely by AI agents are a rising trend, and I expect it’ll continue:

AI-only PR reviews are up. Source: Linear (opens in new tab)
Quality and reliability down
We see this in many of the digital products we use. As Mario Zechner, the creator of Pi, said on the podcast (opens in new tab):
“Everything is broken. It sure feels like software has become a brittle mess, with 98% uptime becoming the norm instead of the exception, including for big services. And user interfaces have the weirdest bugs that you’d think a QA team would catch. I give you that that’s been the case for longer than agents exist. But we seem to be accelerating.”
We covered the phenomenon of quality being worse with more AI usage in March, in ‘Are AI agents actually slowing us down? (opens in new tab)’
Infra capacity shortage
It’s common knowledge that there’s a GPU shortage and memory shortage (opens in new tab) in the market. In addition, there’s also the growing trend of CPU shortages (opens in new tab). As I covered last month:
“Apparently, it’s now nearly impossible to get CPUs on spot instances without long-running connections with cloud providers. Also, reserving specific CPUs now needs to be done months in advance, and cloud providers will even turn down certain reservations because they don’t have enough CPUs or the right type of CPUs.”
If you’re at a company with a non-trivial amount of CPU usage, securing capacity is something worth doing now, as I wrote recently: (opens in new tab)
“The best time to secure more CPU capacity is most certainly right now. I’m hearing rumors that certain cloud regions no longer accept new tenants because all CPU capacity is leased, or negotiations elsewhere are difficult. I’m also hearing that customers are already paying today to reserve capacity that will only come online in data centers from December. This seems predatory by providers, but demand is so high that this is how they likely prioritize new capacity allocation – while earning much higher profits than usual.
If your company has dynamic workloads, and you’ve used spot instances in the past, now could be a good time to allocate fixed capacity – even if it’s more expensive. If you expect meaningful growth, doing so now might mean having options at some cloud providers or in some regions.”
A hit on personal focus and productivity
Software engineer Dima Zaytsev (currently at Linear, formerly my colleague at Uber) told me something that felt very relatable on AI and productivity:
“There’s lots of context switching: there’s always another agent that waits on your response.
Also, it’s just more work. When AI was ramping up, I felt more productive than my peers: I did in an hour what others needed a day for. Now, the expectation has become that everyone works on parallel things. And so, I end up feeling I have more work compared to pre-AI.”
Engineering leaders taking career breaks
An interesting, unexpected trend that has emerged is CTOs, Heads of Engineering, and VPs of Engineering resigning from their jobs, often with nothing specific lined up. From a current CTO who asked to remain anonymous:
“This change [with AI] makes builders want to build. We want to build again, so we can build more, and build faster with better tools, so we can increase the “talent density” of the team by doing more with fewer great people, and so on.
It’s an amazing time to build things. And it’s only getting better. And people just want to be close to it. They want to build and see how far they can go with these new tools.
I have a bunch of friends in my social circle who are the “CTO drop outs” you speak of. More still who are moving from management to IC roles in large companies. And their stories are largely what I describe above. They simply see this as a truly exciting time to build.
In the end, building is why they got into this field in the first place.”
We cover this in the deepdive ‘Headed for the exit: the great engineering leader career break (opens in new tab)’.
4. What’s still the same?
In the LDX3 keynote, I also covered areas that are mostly unchanged from before AI:
Teams are still important. As Katelyn Lesse, Head of Engineering for Claude Platform told me (opens in new tab) in July:
“One thing I’ve heard from some people is “we have two humans and a bunch of agents.” I reply that this isn’t where we’re at. I still have teams whose job it is to own a piece of software, iterate on it, own oncall, and so on. While each of these humans are supercharged by AI, the size and shape of the team is still similar. We still have two-pizza teams.”
Katelyn works at one of the most “AI-pilled” companies globally, so if teams are still as important there as they were before AI, then it’s not a stretch to say that the team structure is still relevant at companies across the industry.
Planning still happens – at least for complex work. Complex projects still have a lengthy planning phase. Also from Katelyn (opens in new tab), on the planning phase of Claude Managed Agents:
“Our planning process looked more like a typical pre-AI planning process. You know how every team has the project, where everyone comes up with some version of the same idea and people keep floating and circling it around until you finally do it? Managed Agents was this for our team. When we started the project, we had documents dating back up to two years about ideas and suggestions.”
Tests and validation are still very important. From Jarred Sumner, creator of Bun, who’s currently at Anthropic:
“The AI writes pretty much all the code, but we have the AI write pretty much all the tests as well. Before AI, for production-ready software, we spent about the same time writing tests as we did on writing code. This is still the case with AI. You need to have a way to trust your code, and tests are probably the best way to.”
At each company I talked to in advance of the LDX3 keynote, I observed that a lot of focus is going into validating the output of AI agents. This makes sense: code review no longer works like it did before, we have more code, and we need to make sure that it won’t break production!
Non-engineers are still not shipping prod code. In the last few weeks, there’s been stuff on social media about AI enabling product managers/designers/non-technical people to ship production code. So, I asked the AI labs, startups, and other companies.
I can report that I did not find any single company where PMs/designers/non-technical people ship to production!
What I did learn about are cases where these folks create bugfixes – often unknowingly! – or new features with their agents that end up with devs to review. There’s also a massive amount of prototyping. But it’s still engineers who are responsible for deciding what can and what cannot be released to production.
We’re rediscovering old patterns that work great with AI. This came up in our podcast (opens in new tab) with Matt Pocock. Matt noticed that agents try to build software layer by layer, which causes bugs between the layers. Reading The Pragmatic Programmer, he discovered the concept of the “tracer bullet.” When he instructed the agent to use “tracer bullets” to build an app (aka implement a “golden path”), the agent started to produce better code.
As mentioned in the podcast, Matt is now reading classic software engineering books to find other “leading words” that guide agents efficiently.
In general, I’m noticing that “old” best practices are helpful when building better software when working with AI. This includes writing unit tests, building a “golden path” first (aka a “tracer bullet”), architecting an application upfront, and even using design patterns. The amusing thing is that we’re talking about decades-old best practices here.
5. What comes next?
It’s possible to look ahead at where the tech industry is headed by identifying trends that are underway, and which look certain to conitnue:
Cloud coding agents + harnesses will dominate. Most devs at companies will run AI agents in the cloud, instead of locally. Ramp offers a blueprint for this in our deepdive about their cloud agent, called Inspect. (opens in new tab) It seems that inside innovative tech companies, the “build your own cloud agent harness” principle is trending.
I also expect vendors like Anthropic, OpenAI, SpaceX, and others to start pushing their cloud coding agent offerings, and for more startups to build their own cloud harnesses.
As engineers, we’ll stop reading the code. It’s too early to tell when this will happen; this year, next year, or further in the future, but it’s likely that few of us will read the code that AI agents produce properly, going forward.
As a rule, I pay a lot of attention to Honeycomb cofounder and CTO, Charity Majors, due to her being a “default sceptic” about new technology, as well as a standout engineer and someone who speaks her mind. On our podcast in August (opens in new tab), she asked:
“What would it take for you to be comfortable shipping code without you reading it and understanding it? Because that is engineering.”
Charity made the point that Ops and QA have already had decades to figure out how to ship code to production that they didn’t write and don’t understand – all while making sure it works! It seems unavoidable that most software engineers will also join this group. This is ironic, since software engineering, for the longest time, was about engineers writing and understanding the code!
Companies will build brand new types of internal infra. We’ll see a “reinvention” of these systems to work well with agents:
CI/CD, with agentic evals part of the pipelines
Agentic software factories, and figuring out what is and isn’t practical
Agents becoming part of deployments, observability, incident management
Agentic observability will surely become its own problem space and specialization: how can we make sure the customer-facing agents deployed work as expected, and how to ensure that systems which agents build and modify actually work as expected?
A “golden age” of refactoring/migrations/full-on rewrites. Migrations and rewrites that had been delayed for years because they’d take similar amounts of time to complete, should take no longer than weeks. That means there’s little excuse to delay any longer!
Engineers understand more about how capable AI agents are with rewrites/refactors/migrations today, and we’ll also clean up tech debt much faster, while having no reason to have a system in a state that we’re unhappy with!
AI “fluency” and “positivity” matter in recruitment at startups. A trend I have not written much about – but which is already happening – is startups selecting engineers for jobs who demonstrate a positive attitude about AI. As the Director of Engineering at a Series D startup told me:
“AI positivity was something we started to screen for in our hiring process. We want people to join who want to help us push the limits with what we can build with AI.”
The reality is that AI coding agents are already everywhere, and agentic systems look set to proliferate very soon. This will bring great demand for software engineers who are willing and able to build next-generation systems, and it’ll be a baseline criterion to be excited about the problem space.
This is not too dissimilar to the way that startups hire engineers who buy into the idea of growing fast, or when companies favor engineers who pick up the tech stack already in use, instead of insisting upon the one that they personally use.
Engineers with deep domain knowledge will become more in demand. In New York, I had a conversation with Titus Winters, lead author of Software Engineering at Google. He told me this observation:
“What you need to succeed anywhere is intelligence, wisdom, and charisma. Intelligence means: how to do it. Wisdom means: what to do. Charisma means: convince others to do it.
When intelligence becomes commonplace, wisdom and charisma become much more important.”
In the context of software engineering, “wisdom” is most easily earned by becoming a domain expert. So, if you work at a fintech company, learn about the finance industry and its customers in order to become a more valuable, in-demand engineer! That also applies if you’re at an agriculture tech company, or work in any other domain.
Finally: remind yourself why you got into the tech industry. At the end of my podcast with Peter Mattis (opens in new tab), the cofounder and CTO of Cockroach Labs, we talked about how it feels pretty exhausting right now to keep up with the pace of change. He said:
“Remind yourself why you got into this in the first place. I got into software engineering because I like building stuff. Now I can build faster, and without some of the compromises that I had before!”
- AI use is shifting from assistance to delegated work: engineers interviewed describe running roughly 5–10 agent sessions in parallel, and the author says nearly all highly productive engineers he met who had hand-written code a year earlier had stopped doing so. Linear reports agent-created issues outnumbering human-created ones since July; GitHub data is described as showing rapid growth in agent-authored PRs, which the author says exceeded human-authored PRs in August. Factory AI reports that 83% of its users use agent skills, up 2.5x from February.
- Internal agent infrastructure is becoming common: the author says most mid-sized-and-larger companies have built their own harnesses, citing examples including Stripe, Uber, Shopify, Google, Meta, and Amazon. Development increasingly starts in Slack and triggers cloud agents, while agentic software factories connect agents to systems such as CI/CD, code review, and deployment; the author expects cloud coding agents and harnesses to expand, including offerings from vendors such as Anthropic and OpenAI and new startup products.
- AI software economics have two opposing signals: companies are adopting open models and smart routing to reduce token costs by 50% or more; at Uber, token use rose while costs stayed flat from May. Meanwhile, the author reports GPU and memory shortages alongside tighter CPU supply, with spot access difficult and some CPU reservations needed months ahead.
- More generated code is straining quality controls: an engineer at a mid-sized startup describes human reviews becoming “LGTM” rubber stamps under the volume, and Linear data shows AI-only PR reviews increasing. The author identifies agentic evaluation in CI/CD and agentic observability across deployed systems and incident management as emerging infrastructure needs. As a reliability caveat, Pi creator Mario Zechner says software feels increasingly brittle, while acknowledging that some of the problems predate agents.
- Leaner project staffing is emerging, but not the disappearance of engineering teams: Anthropic’s Head of Engineering says projects often have no more than two engineers because each is already managing agents, and the author says startup teams are shrinking; she also says teams continue to own software and on-call responsibilities and remain similar in shape to two-pizza teams. One Series D startup says it screens for AI-positive engineers. The author found no company where non-engineers were shipping production code; PMs and designers were prototyping or sending changes to engineers, who retained release responsibility.