
You can be a 100x engineer, today
Let’s start with a couple observations:
-
Grok Bot is a good product, used by hundreds of thousands of users
-
Grok Bot was built in less than 2 months by a small team of engineers running massively parallel agents
Lauren’s latest walkthrough details how she shipped 2k PRs last month. And we know it’s not slop – Grok Bot is a production product with production consequences. Conservatively let’s say other folks on her team only ship 600 PRs/month – that’s still 20 PRs per calendar day. In other words, way too much code to review by hand, let alone write by hand.
And yes, I was an engineer too, and know that PRs can be a vanity metric – that being said, 2k PRs merged is high enough above variance that it indicates a paradigm shift – the sheer volume of ~65 PRs/dev/day indicates that codebases are growing faster than ever before and with less direct human input than ever before.
I’ve heard rumblings of the same from friends at OAI & Anthropic. Code is slung at high velocity. Human code review is a vestige of the past. And while posts have been made here and there on the engineering practices that allow this to happen, I especially thank Lauren (@poteto) for posting so publicly about it and giving us a peek.
Agentic coding can be frustrating. One PR breaks another, one agent stops halfway waiting for your input, another completely misunderstands your request, and sometimes one model just feels dumb on a given day. Yet, in the face of all that, the labs are still pumping out hundreds of PRs a day. So I decided to take a look into how exactly that’s being done.
The answer, like most things, is surprisingly straightforward. Massively parallel agentic engineering boils down to 2 things:
-
Agents that work like humans
-
A repo designed for agents to write error-free code and verify their changes
These two things are all the leverage you need to multiply your impact to that 10x, 100x, or even 1000x engineer. And they’re pretty simple, so I’ll write a little on each below.
#Agents that work like humans
Think about how our coding agents work today. If you use Claude Code or Codex, you issue a prompt, and then based on the thinking level you set, the agent gets to an answer within some rough expected turn length. Then it runs some validation it’s expected to do, perhaps runs tests, and follows some instructions in its AGENTS.md. Then it messages you: “Done!”
Compare this to how humans work. I’ll read code just like the agents do. But I probably won’t just pick an approach based on the code I’m reading and start working right away. I’ll look at the code and try to understand how the entire module works. I’ll look at git history and try to understand why code was written in the way it was or why certain decisions were made. I’ll try to understand how the code interplays with other systems. If the feature is complicated, I’ll come up with a few contending design choices, and ask some teammates to help me pick between them. I’ll implement the change, make sure performance isn’t negatively affected, run a preview build to make sure everything works, put up the PR, and ask a teammate to tear it apart.
This is exactly the workflow that Lauren’s pstack encodes for her agents. The agents will use commands like /how to understand how something works, or /why to dig into the reasoning behind it. They’ll use /architect and /arena to weigh multiple design options, and /interrogate to adversarially review the changes. The list goes on-and-on; point being, Lauren has brought the same engineering rigor she works with to pstack so her agents act the same way.
I was skeptical at first – do my really smart frontier agents really need all these skills? Was it going to make them dumber, or worse? Aren’t we trying to reduce skill bloat, not increase it? Her X post series details how pstack works, so I’ll fast-forward to its impact.
Its effect has been an incredible increase in trust in my agents. No longer do I have to wonder if one agent is going to break a previous existing feature, or if it’s going to take a shortcut without considering all of its downstream implications. I know my agents are going to research deeply into why existing code is the way that it is. I know they’re going to plan out multiple approaches and pick the best one. I know they’re going to request adversarial review on their own code. That means when I get a message from an agent saying it’s done, or it has already merged the PR, I read it with incredibly high confidence: I know the code is right.
I know what you’re thinking here: doesn’t this just consume a ton of tokens? And honestly, the answer is yes. There’s a balance to be struck here and an equation to be weighed. Using pstack takes more time and more tokens – but I’m also able to get a lot more done in a given day, with a lot less babysitting, and a lot fewer errors. So what’s more valuable to me: my token cost, or my time? pstack also allows delegating different responsibilities to different models and different effort levels; so, not everything needs to be on Astra High. A mix of lower effort levels and less capable/expensive models can still yield good results.
I also get some bonus benefits. I'm maintaining my own open-pstack (brings pstack support to Codex & Claude, with support for Grok/Cursor/Devin/OpenCode/Antigravity subagents), which allows me to extend that confidence in the model-choice direction. My distribution supports cross-provider model assignments: meaning I’m able to have SWE-2 write code, have Opus & Sol do investigations + adversarial review, and have Astra plan complex architecture. (This also lets me take advantage of fun promos, like Devin providing free SWE-2 use in Desktop and CLI through October 16, 2026 with any paid plan).
It also helps me avoid “wrong model” anxiety. Every day on X we hear about “this model seems nerfed today” or “this model should only be used on this thinking level” or “this model should only be used for this”. With open-pstack, I don’t really think about those things. The rigor and multi-model nature of open-pstack’s approach to its work basically absolves me of “what if this model is having an off day today”. Which is, funnily enough, analogous to how we human developers used to cover for each other's off days in the past.
#A repo designed for agents to write error-free code and verify their changes
Great, we now have the confidence in tens or hundreds of parallel agents implementing changes to our repo. If that’s the case, then where does that leave the humans? Luckily, we’re able to move our minds onto more high leverage jobs: system architecture & repo gardening.
The former is self-explanatory, but let’s talk about the latter. What is repo gardening? Lauren describes it herself in this tweet – the only thing I’d change is to say that everyone should be a gardener, not just one:
every team needs a gardener. someone quietly watching the stream of PRs flowing into your codebase, noticing the smells: the third isRecord this week, the lint suppressions creeping like ivy across your carefully planned garden. a steady hand tending the weeds that would engulf it in slop if left unchecked
Repo gardening is the process of setting up your project so that agents cannot and will not ship bad code. Some of the gardening is traditional: linting, CI, rules, bugbots, preview deployments. Some of the gardening is pattern recognition: do agents make the same mistakes again and again, and can you come up with skills or rules against it? (An example, Cursor straight up banned comments in their codebase, as their agents were using it as an excuse not to fully implement things).
Much of the most impactful gardening is the work that’s done on the repo to make it such that agents cannot create and propagate bad patterns across the codebase, which agents love to do. So you can weed them out one-by-one, or design the codebase in a way that those patterns cannot materialize at all. Luckily, this comes back to the same engineering skills we developed in the years leading up to agentic engineering; we’ve been designing robust and hard-to-break codebases for humans too. We can ask the same questions of agents that we used to ask of humans. Are agents creating multiple ways to do the same thing? Do we have bad abstractions that are causing our agents a lot of pain? Where is tech debt accumulating, and how can we address it?
A simple example that we’ve probably all seen before: Say your codebase has three different ways to get the current user: some code reads directly from the session, some decodes a token, and some calls a helper. An agent building the next feature copies whichever pattern it encounters first, spreading the inconsistency. You can keep correcting individual PRs, or establish one shared interface for accessing identity and add a lint rule against bypassing it. Now every agent building a feature has a clear pattern to follow, and a check that catches deviations.
All of this gardening coalesces to: how can we create rails such that agents can’t contribute bad code and create a mess, and instead grow the garden in an organized and uniform manner?
When you do this, and combine it with agents that work like humans as described above, you’ve made it possible to contribute 100x more to your repository than ever before. Your impact can finally move as fast as you can think.
That is the job of the 100x software engineer. You’re a farmer, focused on creating the ideal conditions for your crops to grow. Your impact and throughput is multiplied by every agent that works through the rails you’ve set them on – and if you garden correctly, your agents will grow a beautiful codebase that just works.