We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
tsc-rs: an agent-built TypeScript compiler, and what the model choice cost
Theo released tsc-rs (also called ts-rust), a full Rust rewrite of the TypeScript compiler, type checker and LSP. He describes it as an open-source drop-in replacement for tsc (repo) . Agents had been working on it for five months. In his words: "I burned ~$400k of Codex tokens and got nowhere. Burned ~$20k of Opus and got there in 2 weeks." He also says he has not read a single line of the code . He adds that the Opus usage equals about 10 weeks on the $200 Claude subscription .
The port matched its test oracle, upstream TypeScript, closely. Of the first five issues filed, he says four were real upstream TypeScript behaviour that the port reproduced faithfully . He is open about its limits. bun check shipped the same day and is "probably the right choice for most apps using Bun," and tsc-rs is only faster when you use Effect TS checks. He prefers tsc-rs for its drop-in compatibility, WASM readiness and built-in Effect mods .
Mario Zechner and Armin Ronacher explained in a separate conversation why projects like this work. If you have "an Oracle that can tell the LLM if what it did is correct or not, then you basically won." Bun's test suite is enough for agents to drive a port with occasional steering . Ronacher had agents implement every format in a serialization library from sample TOML, CBOR, MessagePack and JSON files, then told them to "fuzz the hell out of it." That found many bugs, and an LLM can reason about cases a fuzz generator misses .
Haiku 5.5: Anthropic's pitch is a cheap subagent
Anthropic says Haiku 5.5 is in Claude Code and the Claude Platform, costs about 75% less to run than Haiku 4.5, and suits use as a subagent under Opus 5.5 or Sonnet 5.5. It suggests "summaries, compactions, or database queries" . Addy Osmani says his team "loves it as a subagent alongside Opus 5.5," and notes it has an adjustable effort setting . Cursor has it under Settings > Models. Pricing is $0.10/$0.50 per million input/output tokens, rising to $0.50/$2.50 above 100k input tokens. Sonnet 5.5 cache reads dropped from $0.20/M to $0.10/M .
Simon Willison found catches that change the routing math:
- Up to 100k tokens it costs exactly the same as GPT-6 Luna. Above that, its price rises 5x, while Luna's goes up only at 272k and only to $0.20/$0.75. For long-context jobs, Luna looks like the better deal .
- The new tokenizer used about 1.25x as many tokens as Haiku 4.5 on the same long prompt, which amounts to a hidden price increase .
-
You can't turn reasoning off. The default is
medium. To try it:llm install -U llm-anthropic,llm anthropic refresh, thenllm -m claude-haiku-5.5 ... -o thinking_effort low. - Max and Team subscribers now get monthly API credits: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team. They don't roll over. You can turn off auto-reload so requests stop when the balance runs out .
ThePrimeagen offers an early counterpoint. Dropping Haiku in for Luna, both with reasoning off, took a one-line config change. Haiku "performed significantly worse" and was slower on both passes and failures, and he doesn't yet know why . Separately, he found OpenAI's new decision model much faster and more accurate than his previous setup at finding click targets in Omarchy QA screenshots . Benchmark your own automation before switching.
Cross-model adversarial review takes one sentence
DHH's technique: when you're driving from Codex, say "Review this with claude", and when you're in Claude, say "Review this with codex". The models know how to start a review through the CLI, take turns and settle an argument. "No magic" . He describes the rest of his setup as "any harness, multiple agents concurrently, barely any skills, and using adversarial reviews" . Theo runs a version of this. Claude spins up Sol subagents and calls Codex through T3 Code to review its work, because he finds OpenAI models "a bit more thorough with their analysis." He still calls the Claude plan "absolutely mandatory" for shipping serious engineering work . On the "nerfed" $200 Codex plan, he argues that cost per task matters more than tokens per dollar. In his own Terminal Bench run, 6.1 Sol at x-high roughly tied Opus 5.5 at about one-thirteenth the cost per task .
Audit your agents for logic they keep reinventing
Theo says his agents wrote over 200 bad "watch PR" scripts. His audit prompt: "audit my history with Claude Code, Codex and other agents on this machine. Look for every time I asked for a PR to be watched or babysat." Then ask how many times the logic was reinvented, how many versions had visible flaws, and roughly how many tokens and dollars were wasted . The production fix in T3 Code took about 30 tries. It replaced CLI calls with direct API calls and switches between GraphQL and REST depending on which uses rate limits more efficiently, with fallbacks. That cut rate-limit usage by over 75% .
Huntley: Nix as the shared environment for agent sandboxes
Geoffrey Huntley argues for a single devenv.nix as the one source of truth for laptops, CI and ephemeral agent sandboxes. His example stanza provides Rust, Postgres, a rustfmt hook and prek "for agent backpressure" . He develops on NixOS and explicitly tells agents to use sudo, counting on rollback. With runNixOSTest, he puts the whole OS, including multi-machine networking and firewall rules, under test . He also uses an overlay to strip force-push out of the Git binary inside agent sandboxes (nix-demo) . The tradeoff he names is that incremental caching is weaker than in Bazel or Buck2 . Separately, he says he is dropping Opus 5.5 because Claude Code's classifier "is too paternalistic and breaks my flow/development loops" .
Smaller items
-
Deep Agents skills: you can now bind tools to a skill with
metadata.include_tools. A bound tool stays out of context until the agent reads that skill . Apps can passpinned_skillsso a skill like/meeting-prepis loaded before the first model call . Settingskills_metadata=Nonemakes the next run rescan the skill library . - Sourcegraph Deep Search lets Claude Code search code you haven't checked out. Their demo finds deprecated packages across Kubernetes .
- OpenAI integrated its desktop app with chat.openai.com in 30 days—an effort Sottiaux said would traditionally take 6–12 months—and he said Astra wrote most of the code.
- OpenAI engineers used Astra to develop the next-generation inference stack and create a version eight times faster in a short time; Sottiaux described engineers collaborating with the model and said improved infrastructure, including Codex, lets teams build faster. This is a model-assisted infrastructure loop that can speed up subsequent development.
- Before a live keynote demo, Sottiaux’s Dot alerted him that production was down, connected the issue to the demo script, and offered to look into a fix; they did not resolve it in time.
- Coding agents work best when correctness has an objective pass/fail check: a large test suite can guide a port even if coverage is incomplete, and humans can steer at checkpoints. For serialization work, giving an agent sample files for formats including CBOR, MessagePack, and JSON, then asking it to fuzz the implementation, exposed bugs; an LLM can also suggest cases a fuzzer may miss.
- Keep humans responsible for issue triage, architecture, and system-level performance: Mario said agents can handle small fixes, but he still reviews reported issues and sees design or architectural work as requiring human judgment. A practical boundary described for Pi was to establish a minimal API first, then have an agent build components against it with rendering tests; agents can help with local performance fixes but may miss broader architectural causes.
- Pi Durable is designed to let long-running agents resume after a process crash without manual continuation or losing tool results, while preserving application state and conversation history. Its task pattern persists an effect’s intent before execution and its result afterward; external effects need idempotency support to recover safely if a crash occurs during execution. Application state can be stored in versioned documents associated with transcript positions, so an agent can recover the state appropriate to a point in the conversation.
- Treat coding agents as parallel workers, with human oversight at module boundaries: Wang says practitioners should be comfortable juggling 5–10 ongoing tasks and using logs, traces, schemas, and input/output data to build evals; he allows implementation slop only in modules whose overall behavior he understands. In his Slack-clone project, agents working at different times created duplicate message paths and an intermittent race condition when context was lost, illustrating the risk of accumulating opaque modules.
- Don’t let AI-generated analysis replace firsthand judgment: Wang favors depth of insight over breadth and describes an employee reporting 14.7% week-on-week video growth without knowing why, after relying on Claude rather than watching the videos; he says people must use their own thinking and domain expertise to assess AI output.
- Evaluate generated software through complete user workflows, not by whether it looks finished: his replacement-app evaluation considered organizer, attendee, sponsor, and speaker perspectives, role-specific logins, and application flows; he says frontier functionality still needs human playtesting. He also describes comparing implementations with point-and-click tests and screen captures against a 200-page flow document.
- Wang gave nontechnical event staff access to modify code with Devon; staff could request changes and get them in roughly one or two hours, and initially skeptical, spreadsheet-oriented staff switched after seeing the submission quality. He contrasts this with a SaaS vendor’s uncertain roadmap timing.
- As a personal, unmeasured estimate, Wang said AI assistance now lets him do the work of roughly four or five copies of himself, noting that work he did with 12 prompts the prior night would otherwise have taken weeks.
- Periodic describes software work moving from GitHub Copilot and early ChatGPT as assistants to Codex as automation improved; it says few of its engineers now write code as before and expects a similar autonomy progression in research tools, potentially shifting from assistance toward pricing outcomes.
- For deciding what to automate, use humans and automation together to find bottlenecks, then prioritize routine tasks that consume substantial staff time and are easy to automate rather than spending months automating tasks where people have a fine-motor-skill advantage. The goal is high-quality, varied data—not full automation for its own sake.
- Periodic builds RL environments from accumulated experimental or computational histories, and can create additional environments around individual tools or tool subsets. It retains conversations, intuitions, lab actions, computations, and code, aiming to train on the process of doing science rather than only its final outputs.
- Periodic uses both open- and closed-source models; it says access to data unavailable to others can improve compute efficiency and, in some cases, let its systems go beyond frontier models. It also cautions that spending substantial compute on noisy data will not produce good results.
- Treat coding-agent work as an iterative design, coding, and verification loop—not a one-shot prompt. Stay engaged in specifying features and UX, checking results, and requesting small changes; distinguish your active steering time from time the agent runs on its own.
- For a CAD prototype, he first used a Fable 5.1 session to explore SDF feasibility and questions such as real-time editing, face extrusion, and holes. He says LLMs write C better, asks for no dependencies, builds a small knowledge archive from relevant papers, and starts with a minimal implementation and tree-based grouping designed to support later operations.
- Make the project testable by the agent: provide ways to create and manipulate objects and capture screenshots; test fillets on random solids from multiple viewpoints and check screenshot consistency. For responsiveness, he uses cached meshes for movement and low-resolution rendering while moving, followed by a GPU render after 500 ms.
- For software intended for AI use, he advocates a well-documented, simple direct API—such as socket commands rather than MCP mediation—and useful outputs for the model, including depth-shaded or wireframe views and measurements.
Sanfilippo used AI as an editor for a story published in Urania—not to rewrite or polish the prose, but to assess its narrative weaknesses—and said the approach worked very well.
- Claude Haiku 5.5 costs $0.10/$0.50 per million input/output tokens through 100,000 tokens, then $0.50/$2.50; its tokenizer uses about 1.25× as many tokens as Haiku 4.5 for the same long prompt. GPT-6 Luna matches Haiku’s price through 100,000 tokens and raises its price only at 272,000, making it a better deal for longer-context workloads.
-
llm-anthropicnow supports new models without a plugin release for each one: update it, refresh the model list, then invoke Haiku throughllm, for example with-o thinking_effort low. Haiku defaults to medium reasoning and does not allow reasoning to be disabled. - Anthropic is adding monthly Claude Platform API credits for Max and Team subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team; credits do not roll over. Auto-reload can be disabled so requests stop when the balance runs out.
- In his own Terminal Bench 4 tests, Theo says 6.1 Soul’s max run scored best with its native harness and its x-high run tied Opus 5.5; he ran the models through Codex and Claude Code and argues that 6.1 Soul’s cost per task was roughly 13× lower in this comparison. His takeaway is to compare completed-task cost under the relevant harness, not just token or subscription value.
- Theo’s cross-model review workflow uses Claude to launch Soul subagents to review its work, and Claude calls Codex through T3 Code to review Claude’s work. He says OpenAI models have been more thorough in his analysis and deeper issue-finding, while he considers the Claude plan the better choice for regularly shipping serious engineering work.
-
Use
devenv.nixas the shared source of truth for toolchains and dependencies across local development, CI/CD, and ephemeral agent sandboxes; Huntley’s example enables Rust, Postgres,prek, and arustfmthook withdevenv up, avoiding separate environment configurations that drift apart. -
Huntley runs coding agents on NixOS and explicitly allows them to use
sudo, relying on the system’s ability to roll back changes; his loop-engineering approach tests the whole OS, usingrunNixOSTestto assert multi-machine networking, firewall rules, and application interoperability before deployment. - In agent sandboxes, he customizes Git with a Nix overlay that removes force-push functionality, then tests the restriction in a NixOS VM; he presents overlays as a way to customize software throughout the dependency stack.
- Trade-off: Nix’s derivation store can complicate incremental caching, an area where Huntley says Bazel and Buck2 excel, though he still considers Nix the best bang for the buck.
- Huntley proposed moving beyond conventional IDEs toward tools designed for collaboration with many LLMs; he says agentic coding products were getting strong results without language-server AST or other symbolic-context injection, so that context may not be essential to agent workflows.
- He cautions that a stronger model alone cannot verify correct production behavior under suboptimal conditions such as flaky networks; agent-built software still needs verification that accounts for runtime conditions.
Reacting to a post naming GPT-6 in Chat, ThePrimeagen said it and “Astra ultrafast” seemed to be creating a “speed addiction epidemic.” The quoted announcement reported a new high of 40M active users across Codex and ChatGPT Work and a banked reset in paid accounts.
- Stacklok’s open-source Mecatl harness, begun in June, is designed to run agents in the cloud: its agent loop is independent of the client, model provider, state store, and execution environment, while tool calls and bash, session management, and memory are separated into components that can be centrally managed rather than kept in local JSONL files.
- Stacklok’s open-source ToolHive manages MCP servers; its Kubernetes-based offering includes a gateway, registry, and operator helper, with input/output authentication support. Stacklok’s AI Gateway handles access control, budgets, reporting, and provider routing, but does not currently select models by task; the company favors semantic model selection within the harness, where more context is available. The gateway was not yet open source and was described as planned.
Theo disputes the $132M/year calculation for a Cursor profile reporting 1.5T tokens: he says it overstates costs by applying an inaccurate rate to usage that includes cached tokens, and estimates API costs of $100K–$500K/month instead. He estimates 1.5T tokens would cost about $800K with expensive models or as little as $100K with cheaper models and more cache hits. He says his own 33.5B-token usage “came out to $0.56/m”; in a separate reply, he reports a blended $0.46 for real-world Opus 5.5 use, nearly 20× below the $8/M assumption. Theo cites a 30× drop in cost at a given intelligence level over five months and projects that a $100K/month workload could fall to $3K/month in a few months and $100/month next year if the trend continues.
For model-generated components or migration plans, separate review into two filters: use experience-based taste to reject weak drafts quickly (potentially most of them), then use judgement to choose what to deliver by weighing quality against time, risk, and user or team costs.
Theo says agents worked on tsc-rs, a Rust rewrite of the TypeScript compiler, type checker, and LSP, for five months. He reports spending about $400k in Codex tokens without getting there, then about $20k on Opus to complete it in two weeks; he says he had not read any of the code. He describes tsc-rs as an open-source drop-in replacement for tsc, available now.
Addy Osmani says his team uses Claude Haiku 5.5 as a subagent alongside Opus 5.5; Haiku 5.5 is about 75% cheaper to run than Haiku 4.5 and has an adjustable effort setting. Claude says it is available in Claude Code and recommends it for high-volume, cost-sensitive subagent tasks such as summaries, compactions, and database queries, alongside Opus 5.5 or Sonnet 5.5.
ThePrimeagen tested Claude Haiku 5.5 against OpenAI GPT-6 Luna, both with reasoning set to none, on an automation task; swapping in Haiku required only a one-line config change, but it performed significantly worse and ran significantly slower on both passes and failures. The reason for the difference was still unclear.
Haiku 5.5 is described as a significant step up from Haiku 4.5 for coding, computer use, and knowledge work . Alex Albert says Haiku 5.5 is much faster and 75% cheaper, noting that Haiku 4.5 came out on October 15, 2025, less than a year before .
ThePrimeagen tested OpenAI’s new decision model on Omarchy QA results, using the same prompt and image, and found it generally faster and more accurate at identifying screenshot targets to click; the post gives no quantitative comparison. In one example, it correctly inferred the terminal background from an image and clicked it.
Geoffrey Huntley says /z80 was productionized as an MCP server and a paper was published; he characterizes it as using an LLM to clone software and intellectual property (“running the LLM like a Bitcoin mixer”).
the world hasn’t figured out yet that you can literally just fix everything with a Nix overlay
There are many things that are uncertain in our industry right now, but one thing I am certain about, and have been for almost 13 years, is that Nix is a terrible programming language.
I remember when I first learned it, it felt like pushing shit uphill. The learning cliff is ferocious, and it took me a couple of years to master it because there was no AI back then, but I stuck with it. There’s some magic here, and in this post, I’m going to show you why Nix should be your primary choice for all your software projects.
First, I’m going to open with this: many pieces of technology are in our heads as really hard to learn, complex, and maybe that’s stopped you from picking one up, but I want to encourage you to falsify those thoughts any time you start thinking along those lines. Now that we have AI, things that used to be advanced power tools, hard to use or designed for masters, are now accessible to everyone. You can just prompt for outcomes.
first, some introductory knowledge
The term Nix is overloaded. It means many different things; when someone says they use Nix, the first thing you should do is ask, “How do you use Nix?” and “What is Nix to you?” because there are many ways to use it; it’s not just a package manager or a build system; it can also be an operating system.
In this post, I’ll focus on Nix’s versatility and utility, and why it’s so powerful in the age of AI. There’s a reason the labs are using this to build the models you are consuming…
agent and developer experience
By using tools like https://devenv.sh/ (opens in new tab), you can define a single source of truth for your compilation toolchain and required third-party dependencies. Humans can use this single source of truth in CI/CD and in ephemeral sandbox development environments used by agents.
# visit https://devenv.sh/getting-started/ for installation instructions
$ vi devenv.nix
# devenv.sh/packages/
packages = [ pkgs.prek ];
# devenv.sh/languages/
languages = {
rust.enable = true;
};
# devenv.sh/services/
services = {
postgres.enable = true;
};
# Configure and install the Rust formatting Git hook.
git-hooks = {
hooks = {
rustfmt.enable = true;
};
};
$ devenv upthis one stanza provides rust, postgres and prek (for agent backpressure) and it just works across all operating systems
I keep coming across clients with drift between how a local laptop is configured and how their CI/CD is configured. And now we’ve got ephemeral sandbox environments; they’re heading down a path of triplicating that drift. This is utter madness. It is not needed. Stop it.
You might be thinking, “Well, there’s mise for this.” Trust me, mise isn’t good enough. Mise is a low-power, low-IQ tool. It does one thing, and it does it well, but it doesn’t enable your agents to truly fly.
not just a package manager
nix is super composable. That same expression that defines your human developer environment setup, your CI/CD setup, and your ephemeral sandbox setup can also be reused to build performant Docker images.
pkgs.dockerTools.buildLayeredImage {
name = "git";
tag = "latest";
contents = [
(pkgs.buildEnv {
name = "image-root";
paths = [ pkgs.git pkgs.cacert ];
pathsToLink = [ "/bin" "/etc" ];
})
];
config.Env = [
"PATH=/bin"
"SSL_CERT_FILE=/etc/ssl/certs/ca-bundle.crt"
];
config.Cmd = [ "/bin/git" "--version" ];
};This expression here creates your docker image with git in it. But there’s no reason why this docker image can’t be the binaries used to run your application in production. 🫡
But perhaps… the real reason I love Nix is its power to compose an operating system. With a standard operating system such as Debian or Ubuntu, what happens if an agent is given sudo access to do things on your machine?
You’d be pretty scared, right?
What if I told you that when I’m doing agentic development, I develop on NixOS and I explicitly prompt my agents to use sudo as part of my loop engineering, and it is safe because NixOS is designed to make it nearly impossible to break a machine, and if it does, you can instantly roll back that change.
why the operating system matters
When others do loop engineering with their agents, they’re likely running loops with building the application, perhaps Postgres, but rarely anything more. The primary difference between how others do loop engineering and how I do is that, when I run loops, I put the entire system under test, and the entire System Under Test IS the operating system.
You might think, “Whoa, what the —“ Like, how the hell do you do this?
ah! My sweet summer child. It is really simple.
NixOS has a testing framework (runNixOSTest), built in, and for anything you could ever want to assert about how an operating system is configured, or even across many machines, you can spin up a cluster. You can use a number of machines in a test, assert the network rules between them are correct, and that the right IP tables are forwarded or dropped. You do machines of machines, and you actually test the interoperability of that environment against your application.
Here’s an example of what this looks like when you’ve defined an operating system in NixOS and want to assert the network is configured correctly with your firewall zones, and that network activity works for your application before you deploy it.
# NixOS VM test: HAProxy in front of a Python hello-world server.
#
# Topology (two QEMU nodes, separate L2 networks):
#
# vlan 1 192.168.1.0/24 frontend
# haproxy eth1 192.168.1.1/24 HAProxy binds :80 here only
#
# vlan 2 192.168.2.0/24 backend
# haproxy eth2 192.168.2.1/24 allowed source
# eth2 192.168.2.50/24 extra address, must be rejected
# web eth2 192.168.2.2/24 Python hello-world on :8080
#
# NixOS test IP scheme is 192.168.<vlan>.<nodeNumber>. nodeNumber comes from
# the sorted node name, so haproxy = 1 and web = 2.
#
# web firewall:
# - inbound: accept TCP 8080 only from 192.168.2.1, drop everything else
# - outbound: no new connections; established replies to HAProxy still pass
#
# Run:
# nix build .#checks.x86_64-linux.haproxy-hello
# nix run .#driverInteractive # drop into the test driver
{ lib, ... }:
let
frontendVlan = 1;
backendVlan = 2;
# Sorted node names: haproxy, web.
haproxyNode = 1;
webNode = 2;
haproxyFrontendIp = "192.168.${toString frontendVlan}.${toString haproxyNode}";
haproxyBackendIp = "192.168.${toString backendVlan}.${toString haproxyNode}";
# Same backend VLAN, but not the address the firewall allows.
haproxyUntrustedIp = "192.168.${toString backendVlan}.50";
webIp = "192.168.${toString backendVlan}.${toString webNode}";
frontendPort = 80;
webPort = 8080;
helloBody = "hello from python";
in
{
name = "haproxy-hello";
nodes = {
haproxy = { pkgs, ... }: {
virtualisation.vlans = [ frontendVlan backendVlan ];
virtualisation.memorySize = 512;
environment.systemPackages = [ pkgs.curl ];
# Harness already assigns 192.168.2.1/24. This alias is the negative
# control: same L2 network, source the backend must refuse.
networking.interfaces.eth${toString backendVlan}.ipv4.addresses = [
{
address = haproxyUntrustedIp;
prefixLength = 24;
}
];
networking.firewall = {
enable = true;
# Frontend lives only on vlan 1. Do not publish :80 on the backend net.
interfaces.eth${toString frontendVlan}.allowedTCPPorts = [ frontendPort ];
};
# The haproxy module already opens a \`global\` section and installs the
# stats socket, so this value must start at the next section.
services.haproxy = {
enable = true;
config = ''
defaults
mode http
timeout connect 5s
timeout client 10s
timeout server 10s
timeout check 5s
frontend fe_http
bind ${haproxyFrontendIp}:${toString frontendPort}
default_backend be_hello
backend be_hello
option httpchk GET /
http-check expect status 200
server hello ${webIp}:${toString webPort} check
'';
};
};
web = { pkgs, ... }: {
virtualisation.vlans = [ backendVlan ];
virtualisation.memorySize = 512;
environment.systemPackages = [ pkgs.curl ];
networking.firewall = {
enable = true;
# extraCommands is the iptables backend. Pin it so the rules below apply.
backend = "iptables";
allowPing = false;
logRefusedPackets = true;
# Nothing open on every source. The only accept is extraCommands.
allowedTCPPorts = [ ];
extraCommands = ''
iptables -A nixos-fw -p tcp -s ${haproxyBackendIp} --dport ${toString webPort} -j nixos-fw-accept
# No new outbound. Established replies (the HAProxy request path) still pass.
iptables -A OUTPUT -o lo -j ACCEPT
iptables -A OUTPUT -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT
iptables -A OUTPUT -j REJECT --reject-with icmp-host-prohibited
ip6tables -A OUTPUT -o lo -j ACCEPT
ip6tables -A OUTPUT -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT
ip6tables -A OUTPUT -j REJECT --reject-with icmp6-adm-prohibited
'';
extraStopCommands = ''
iptables -D nixos-fw -p tcp -s ${haproxyBackendIp} --dport ${toString webPort} -j nixos-fw-accept || true
iptables -D OUTPUT -o lo -j ACCEPT || true
iptables -D OUTPUT -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT || true
iptables -D OUTPUT -j REJECT --reject-with icmp-host-prohibited || true
ip6tables -D OUTPUT -o lo -j ACCEPT || true
ip6tables -D OUTPUT -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT || true
ip6tables -D OUTPUT -j REJECT --reject-with icmp6-adm-prohibited || true
'';
};
systemd.services.hello-world = {
description = "Python hello-world HTTP server";
wantedBy = [ "multi-user.target" ];
after = [ "network.target" ];
serviceConfig = {
ExecStart = "${pkgs.python3}/bin/python3 ${pkgs.writeText "hello.py" ''
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
BODY = b"${helloBody}\n"
class Handler(BaseHTTPRequestHandler):
def do_GET(self):
self.send_response(200)
self.send_header("Content-Type", "text/plain")
self.send_header("Content-Length", str(len(BODY)))
self.end_headers()
self.wfile.write(BODY)
def log_message(self, fmt, *args):
return
if __name__ == "__main__":
ThreadingHTTPServer(("${webIp}", ${toString webPort}), Handler).serve_forever()
''}";
DynamicUser = true;
NoNewPrivileges = true;
ProtectSystem = "strict";
ProtectHome = true;
Restart = "on-failure";
};
};
};
};
testScript = ''
start_all()
haproxy.wait_for_unit("haproxy.service")
web.wait_for_unit("hello-world.service")
web.wait_for_unit("firewall.service")
haproxy.succeed("ip -4 addr show dev eth${toString frontendVlan} | grep -q '${haproxyFrontendIp}/24'")
haproxy.succeed("ip -4 addr show dev eth${toString backendVlan} | grep -q '${haproxyBackendIp}/24'")
haproxy.succeed("ip -4 addr show dev eth${toString backendVlan} | grep -q '${haproxyUntrustedIp}/24'")
web.succeed("ip -4 addr show dev eth${toString backendVlan} | grep -q '${webIp}/24'")
with subtest("backend accepts only haproxy on the hello port"):
haproxy.wait_until_succeeds(
"curl -fsS --interface ${haproxyBackendIp} http://${webIp}:${toString webPort}/ | grep -F '${helloBody}'"
)
haproxy.fail(
"curl -fsS --max-time 5 --interface ${haproxyUntrustedIp} http://${webIp}:${toString webPort}/"
)
haproxy.fail(
"curl -fsS --max-time 5 --interface ${haproxyBackendIp} http://${webIp}:9/"
)
with subtest("haproxy on the frontend vlan proxies to python"):
haproxy.wait_until_succeeds(
"curl -fsS http://${haproxyFrontendIp}/ | grep -F '${helloBody}'"
)
with subtest("python host cannot open new outbound connections"):
web.fail("ping -c 1 -W 3 ${haproxyBackendIp}")
web.fail("curl -fsS --max-time 5 http://${haproxyBackendIp}:${toString frontendPort}/")
# Return path still works: the proxy subtest above already required it.
web.succeed("iptables -S OUTPUT | grep -q 'REJECT'")
web.succeed(
"iptables -S nixos-fw | grep -q -- '-s ${haproxyBackendIp}/32 -p tcp -m tcp --dport ${toString webPort} -j nixos-fw-accept'"
)
'';
}There’s an old saying that bare metal was complicated and it was a mess. This is no longer true now that we have Nix. Honestly, one of the most galaxy-brain moves you can make these days is to ditch hyperscalers like AWS and acquire bare metal to build that fleet with NixOS, because business margins are going to get compressed by AI.
Sure, you’ll need to find an older, more experienced sysadmin, but once you get the right patterns in place, one or two of those people can operate with leverage equivalent to a team of 50 “cloud certified” monkeys. I shit you not, no exaggeration, it is insane what you can do with this composability.
single sources of truth
Nix, when architected correctly, enables you to define a single source of truth for any piece of software in your stack, including dependencies.
Let’s say there is a security vulnerability, and there is, like, a critical OpenSSL vulnerability. How long would it take you to identify all the different OpenSSL versions deployed within your organization? How long would it take to patch, including all dependencies of third-party software that depends on that particular version within your organization, including software dependencies linked against it?
With nix, this can be achieved through a couple of lines
{ pkgs, ... }: {
nixpkgs.overlays = [
(final: prev: {
openssl = prev.openssl.overrideAttrs (old: {
version = "3.0.7";
src = prev.fetchurl {
url = "https://www.openssl.org/source/openssl-3.0.7.tar.gz";
hash = "sha256-...";
};
}what you just saw was an overlay
An overlay is the key primitive which makes Nix absolutely amazing. I want you to think about how you use open source software. What happens when you find a bug? Do you go to GitHub and ask someone to take your pull request? Or maybe you fork it. But when you fork it, you have to figure out how you’ll build it and where you’ll host the resulting artifact (ugh, Artifactory) so it can be consumed.

I’ll show you a concrete example. It’s here on GitHub. It shows a pattern I’ve used for years: in my agent sandbox environments, I customize the Git binary and remove the agent’s ability to force-push by removing that functionality from Git itself. It also includes all the patterns shared in this blog post, from building a Docker image to building a NixOS virtual machine and testing that git push --force has been removed.
What happens if what you need to customize isn’t just a standard command-line tool? What happens if it’s a low-level thing, like OpenSSL, that many things depend on, and you need to rebuild the world? Think about that graph. Think about the pain.
With Nix, every change is just an overlay; all software becomes infinitely customizable at all levels, all the way down to the Linux kernel itself, and you can patch anything by asking an agent to customize it. The only restriction on what you can do is your imagination and ambition. These LLMs know Nix very well because the labs themselves are using Nix on their journey toward reaching RSI…
ps.
You might be wondering about Bazel and Buck2. I also use these alongside Nix, but the best bang for the buck right now is just to use Nix. There are a lot of problems with Nix, for example, how the derivation store works and what that means for incremental caching, something Bazel and Buck2 excel at, but I’m not going to complicate things right now by going into these details. Move past this and just start using Nix to customize all your software.
pps.
If you are building security-critical systems, burn tokens to find every security vulnerability in third-party dependencies your software stack relies on using a cyber model, and fix them with overlays. 🧠🧠🧠
pps. socials
🗞️ the world hasn’t figured out yet that you can literally just fix everything with a Nix overlay
In this blog post, I go into how I use Nix, why you should use it, and in the tweet below is a runnable demo showing all of the concepts in this post https://t.co/rZ3sqCb5gt (opens in new tab) pic.twitter.com/QeyAHWTxnR (opens in new tab)
— geoff (@GeoffreyHuntley) October 7, 2026 (opens in new tab)
-
Use
devenv.nixas the shared source of truth for toolchains and dependencies across local development, CI/CD, and ephemeral agent sandboxes; Huntley’s example enables Rust, Postgres,prek, and arustfmthook withdevenv up, avoiding separate environment configurations that drift apart. -
Huntley runs coding agents on NixOS and explicitly allows them to use
sudo, relying on the system’s ability to roll back changes; his loop-engineering approach tests the whole OS, usingrunNixOSTestto assert multi-machine networking, firewall rules, and application interoperability before deployment. - In agent sandboxes, he customizes Git with a Nix overlay that removes force-push functionality, then tests the restriction in a NixOS VM; he presents overlays as a way to customize software throughout the dependency stack.
- Trade-off: Nix’s derivation store can complicate incremental caching, an area where Huntley says Bazel and Buck2 excel, though he still considers Nix the best bang for the buck.