Pi is an open source terminal coding agent from Earendil, and on 2026-10-01 it reached 1.0. Two posts went up the same day, the 1.0 release and a new package called Pi Durable, and both landed on the Hacker News front page, with 1680 and 505 points. By 2026-10-05 the GitHub repo (https://github.com/earendil-works/pi) showed 112.5k stars and 14.3k forks. This post explains what shipped, how the most interesting part of it works, and why it matters if you build web apps with AI.
What happened
Earendil describes Pi 1.0 as "a hardened, minimal, extensible agent harness" (https://earendil.com/posts/pi-1-0/). A harness is the program around the model: it holds the conversation, calls tools, edits files and decides what goes into the prompt. Pi is MIT licensed, and Earendil says hundreds of thousands of people use it every week. They give no exact figure, so that claim is theirs alone.
The 1.0 feature list:
- Codemode, with native support for MCP servers, non-LLM models such as Jev, and image models.
- Virtual models, which extensions can define.
- Deferred tool loading, so tool definitions are not all loaded up front.
- Cache warming for Anthropic models.
- Mid-conversation system messages.
- A full-screen terminal UI by default, with a new theme.
The release notes (https://github.com/earendil-works/pi/releases) add detail. Codemode first landed in v0.99.0 on 2026-09-29, together with MCP server integration. In 1.0, Codemode uses "about 40% fewer prompt tokens" and recovers from errors better, scripts can generate images, and MCP OAuth was hardened to RFC 9207. Two patch releases followed fast: v1.0.1 on 2026-10-03 added project-level MCP server overrides through a .pi/mcp.json file and a Nix flake install, and v1.0.2 on 2026-10-04 added per-thinking-level sampling settings for OpenAI-compatible APIs.
Alongside it came Pi Durable (https://earendil.com/posts/pi-durable/), which Earendil labels "experimental, and the API might still change". It is a framework for long-running agents that survive a crash and can be shared by several people at once.
The same day, a separate story reached Hacker News at 187 points: "Figma restricts MCP access to whitelisted clients, excluding Pi". I covered that allowlist in the last post, so here it is only context: Pi bet on MCP in the same week a major MCP server closed its door to it.
How Codemode works
The problem Codemode attacks is tool bloat. The usual way an agent uses MCP is to load every tool's name, description and JSON schema into the prompt, so the model can pick one. Connect a few servers and the prompt carries dozens of schemas on every turn, whether the task needs them or not. That costs tokens, and it costs attention: the model has more to read before it does anything.
Codemode turns this around. Instead of the model calling tools one at a time from a list in its context, it writes a short script, and the script calls the MCP tools it needs when it runs. The model only has to know that the tools exist and how to reach them, not hold every schema at once. A chain of five calls becomes one script with a loop and some conditions, rather than five round trips through the model.
That is where the "about 40% fewer prompt tokens" comes from, though the release notes give no base figure and no model, so treat it as a direction rather than a benchmark. The same idea explains why non-LLM models fit in: a script can call Jev or an image model the same way it calls an MCP tool.
Deferred tool loading is the other half. Even without Codemode, Pi can hold back tool definitions until they are needed. Claude Code does something similar for its own tools, which tells you this is now a shared concern across harnesses, not one vendor's trick.
How Pi Durable works
Pi Durable is about what happens when an agent runs for a long time. Each task saves a checkpoint before it moves on. If the process crashes, the task resumes from the last checkpoint instead of starting over. State can live in memory, in SQLite or in JSONL files. Several clients can watch and steer one conversation, so two people, or a browser tab and a terminal, can follow the same agent.
It is not small: Earendil puts the source at about 15,000 lines, and the repository ships over thirty examples. One of them, a vacation planner, is about 1,300 lines of TypeScript. That example is the useful signal. It says Pi Durable is meant for any agentic app, not just coding agents.
Why someone building web apps with AI should care
Three reasons, in order of how soon they matter.
You can own the harness. Pi's extensions, skills, prompt templates and themes ship as npm or git packages, and a TypeScript SDK embeds Pi in an app. If your product needs an agent, you can start from a harness you can read and change, under an MIT license, instead of wrapping a closed CLI.
Durable sessions are the hard part of an agent tab. Any web app that adds an "agent" feature runs into the same problems: the user closes the tab, the server restarts, a second person opens the same task. Checkpoints, resumable tasks and several clients on one conversation are exactly that list. Pi Durable is a TypeScript answer to it. It is also experimental, so it is something to prototype on, not to put under a paying customer yet.
Tool bloat is a harness problem. If you are building MCP servers for your own product, Codemode is a reason to keep each tool small and well named. A harness that writes scripts against your tools rewards clear, composable tools more than one giant tool with twenty options.
What is not known yet
- The 40% figure's base. The release notes give the percentage but not the model, the task or the starting token count.
- Figma's position. The original post about the restriction did not load for the research, and there is no word on whether Pi will be added to Figma's list.
- Sandboxing. The 1.0 post does not mention any built-in sandbox. How you isolate an agent that runs scripts on your machine is left to you.
- Pi Durable in production. The API may still change, and no production users are named.
My take
I think Codemode is the right idea, and it is the idea I would copy first. My own daily setup connects a lot of MCP tools, and the real cost is not one expensive call. It is every turn carrying schemas the task never touches. Letting the model write a small script against tools it discovers on demand is closer to how I would want a junior developer to work: look up the API when you need it, do not memorise the whole SDK first.
Where I am more careful is the unknowns. An agent that writes and runs scripts is more powerful than one that calls tools one by one, and with no sandbox mentioned, the isolation is my job. And Pi Durable is the part I would most like to build on, a resumable agent session in TypeScript, but "experimental, and the API might still change" is a clear label. I would try it in a side project this month and wait for a stable API before I put it behind a real product.
