I built the driveshaft first
The first agent I put into production at a large payments company was hardwired into the chat endpoint: its own code branch, its own builder imports, its own page. It demoed beautifully. Then came the second agent, and the third, and by the time I was staring at a websocket handler with a growing if/elif chain deciding which bespoke agent to boot, I understood that I had built a driveshaft. Every new capability had to physically bolt onto the same shaft, and the shaft dictated the shape of everything around it.
The contributor doc for that codebase now contains a rule written in scar tissue: no per-agent branches in the chat runtime, ever. I wrote the rule because I had lived the alternative, and the living is the point of this essay. I originally published a long series on technology transitions, electricity to spreadsheets to the internet to AI. This piece replaces it. History gets four short lessons, each one paired with where I hit it in production, and the rest of the space goes to the place where AI refuses to follow the script.
The pattern, compressed
Every general-purpose technology arrives the same way. It gets judged against the thing it seems to replace, adopted unevenly by people experimenting without permission, governed late, and it pays off only after someone rebuilds the work around it. The tool is the cheap part. The complement, the rewiring of workflows and review and governance, is where the advantage comes from and where the years go.
| Transition | The lag | What actually moved the needle |
|---|---|---|
| Electricity | Motors sat in steam-era factory layouts for years; measured productivity followed decades behind availability | Rebuilding the floor around distributed power |
| Spreadsheets | Adoption was nearly instant; the risk surfaced later, buried in formulas | Model review and audit disciplines that scaled with the new producers |
| Internet | 14 percent of US adults online in 1995, 96 percent today (Pew) | Governance built after the fact: security, privacy, trust infrastructure |
| AI | 53 percent population adoption within three years (Stanford AI Index 2026) | Still being decided. My bet: continuous verification, for reasons below |
Electricity: the rewiring lag, lived
The first electrified factories dropped motors into the spots where steam engines used to sit and left the layout alone, because the layout was the shape of the old power source. Belts and shafts had dictated where machines could stand. The productivity everyone had been promised showed up only when factories were rebuilt around what distributed power made possible: machines placed by the logic of the work. The lag lived in the architecture, not in the motor.
My bespoke agents were motors bolted to a driveshaft. The rewiring, when I finally committed to it, meant converting the product into a manifest-driven orchestration platform: one shared runtime, agents declared as data rather than wired as code, and a compiler that assembles each agent in ten steps, resolving access, tools, permissions, context, prompt, model, and graph strategy per session. On the origin product that came to 15 agents defined as manifests. The general platform that grew out of it now runs 16 registered agents on the identical runtime, and adding one means writing a manifest file and registering it. The marginal cost of an agent collapsed from a bespoke build to roughly a pull request, which is the board-level version of what rewiring buys: the same headcount now ships agents instead of plumbing.
The honest part, because every rewiring story skips it: the conversion took months, and during those months my visible feature velocity was close to zero. A rewiring is a capital project wearing a refactor's clothes. If you budget it as a refactor, you will abandon it halfway and end up with two driveshafts.
Spreadsheets: democratized power, hidden risk, lived
Lotus 1-2-3 pulled in 53 million dollars in its first year because it changed who was allowed to model a business, and Steven Levy saw the trap as early as 1984: the bad assumptions hide in the formulas and relationships, not the visible numbers. A spreadsheet can be immaculate in every cell and wrong in its structure, and the polish tells you nothing.
AI runs the same play with the volume turned up, and I have the receipts from my own QA. When I migrated my analytics agents to a new tool contract, I ran three simulated user personas through two rounds against the semantic layer and checked every claim the agent made against the database itself: 118 assertions. It passed 88, a 74.6 percent pass rate. All 20 hard failures were the same species of failure: fluent numbers that did not match the ground truth. The agent claimed 221 orders where the database held 835. It claimed 5,121 guests where the database held 1,320, a 288 percent inflation, delivered in perfectly confident prose. Nothing in the tone of those answers distinguished them from the 88 correct ones.
The spreadsheet era's answer was review that scales with the new producers, and that is the answer here too, with one upgrade: the reviewer cannot be a human reading transcripts, because fluency defeats human review at volume. My review layer is 76 golden tests that replay real questions against the agents and verify the answers against the database, plus provenance attached to every query. Review did not scale by adding reviewers. It scaled by making the database the reviewer.
The internet: governance lag, lived
In 1995, Pew found 14 percent of US adults online. Clifford Stoll called the internet an ocean of unedited data in Newsweek that year; Robert Metcalfe predicted it would collapse in 1996 and later, literally, ate his words. Today the figure is 96 percent. The skeptics were wrong about adoption and frequently right about governance: the collapse never came, but the spam, fraud, and privacy loss all did, because the institutions arrived decades after the capability.
The lesson is to hold adoption and governance as separate questions, and the production version of that lesson is about where governance physically lives. Early on, my governance was documentation, which is to say it was advisory. Now it is compiled. The platform compiler runs three gates on every session before a model sees a request: an access gate that decides whether this user can reach this agent at all, a permission gate that filters the tool catalog down to what this user is actually allowed to invoke, and a domain filter on top of that. Governance in a policy PDF lags adoption the way internet governance did. Governance in the compile path ships with every request, and nobody has to remember to follow it.
Where AI breaks the pattern
Everything above says AI is a normal general-purpose technology and the old playbook applies: expect the lag, rewire the work, scale the review, govern before scale. I believe all of it. And then AI breaks the pattern in one place that changes what the playbook can promise.
Every previous transition technology held still while you rewired around it. Electricity never lied about its torque; once the factory was rebuilt, the motor's spec sheet stayed true. AI's capability is jagged and it moves. The Harvard Business School field experiment with BCG found consultants using GPT-4 on a task deliberately chosen to sit outside the model's competence were 19 percent less likely to reach a correct answer than colleagues working unaided, while the same tool made them measurably better inside the frontier. My own data shows the same jaggedness at close range: in that 118-assertion QA run, single-period revenue questions landed within 1 to 2 percent of database truth while multi-period questions missed by 46.9 percent. Same agent, same semantic layer, questions a user would consider adjacent.
And the frontier does not just vary across tasks; it shifts under your feet across model generations. I once swapped the frontier model in my autonomous development factory mid-session, same harness, same project state, and wrote up what happened: the failure types changed, the token economics changed, the verification behavior changed, all inside one afternoon. A factory rewired for one motor woke up with a different motor in it.
That is the break. The historical playbook assumes redesign is a phase: you rewire once, then harvest. With AI, the capability you rewired around is neither uniform nor stationary, so the highest-value complement turns out to be a permanent verification layer that re-measures the frontier continuously, ahead of the redesign itself. The golden tests, the provenance, the gates: those are not scaffolding I get to remove when the transition completes. They are the organ that makes the transition survivable, because the transition does not complete.
The stance
History's advice compresses to one sentence: the technology sets the ceiling, and the rewiring decides how close you get. That sentence survived electricity, spreadsheets, and the internet, and it survives AI. What does not survive is the assumption of a finish line.
So here is the claim I expect pushback on. If your AI roadmap has a completion date, you have already mispriced it. You have priced a procurement, and what you are actually buying is a permanent operating capability with a permanent verification cost attached. I would rather fund a smaller transformation with no end date than a larger one with a ribbon-cutting, because I have watched my own agents pass 88 of 118 checks while sounding identical on all 118. The organizations that win this transition will be the ones that never declare it finished.