We've Been Here Before: Mainframes, Calculators, and the AI Correction
I started my career on the tail end of mainframe technology — glass-house computing, dumb terminals, IT as the sole owner of the machine. I watched Netware and Windows Server rise and get sold, explicitly, as liberation from that model: put the power on the desktop, let departments run their own file and print servers, break the priesthood's monopoly on the machine room. I watched the dot-com years bolt a network-facing layer onto that decentralized world, often literally in a server closet down the hall. I watched the on-prem buildout that followed — racks you could walk up to, SANs, blade servers — consolidate all over again before "cloud" was even the word for it. And now I'm watching the industry arrive, with straight faces and trillion-dollar capital budgets, at something that looks a great deal like where I started.
This isn't nostalgia. It's a pattern, and it's worth taking seriously as one, because AI is currently running two versions of it at the same time — one in infrastructure, one in cognition — and the two cycles are colliding in a way that makes this moment different from any previous swing of the pendulum.
The pendulum, not the arrow
Every generation in this story tells itself the same origin myth: this time we're escaping centralization for good. Client-server was sold as freedom from mainframe vendor lock-in. The web was sold as freedom from the corporate LAN. Cloud was sold as freedom from the on-prem data center you had to staff and cool and depreciate yourself. Each time, the freedom was real for a while, and each time, the industry found its way back to something centralized — just owned by a different vendor, wearing a nicer wrapper.
Cloud isn't identical to a mainframe, to be fair to the marketing. Elastic, pay-by-the-hour provisioning has no clean historical analog — a mainframe was sized for peak load and you ate the idle cost the rest of the time. APIs turned infrastructure into something programmable rather than merely rentable. Those are real differences, not cosmetic ones. But the underlying shape — centralized compute, thin clients, someone else's building, a trust boundary you don't control — is the same shape it's always been. The wrapper changes. The physics of who owns the machine and who rents access to it keeps rhyming.
AI is the next turn of this wheel, but it's turning in two directions simultaneously, which the old cycle never quite did.
Splitting down the middle
Training a frontier model requires an amount of tightly-coupled compute that makes even hyperscale cloud look modest — you can't distribute a training run across scattered nodes the way you can distribute web serving, because the interconnect bandwidth between GPUs matters as much as the GPUs themselves. That pulls capital into a handful of sites, more extreme in its concentration than anything client-server or early cloud produced. Mainframe logic, but denser.
Inference is a different animal. Once a model exists, running it doesn't need the same tight-cluster interconnect, which is why on-device inference — phones, laptops with NPUs, smaller distilled models — has real momentum. Latency, privacy, and cost all push inference toward the edge, even as training concentrates harder than ever. So instead of a clean pendulum swing, you get a bifurcation: mainframe-style concentration for making the model, client-server-style distribution for using it.
That split matters, because it means the infrastructure story and the economics story aren't actually the same story — and conflating them is where a lot of the current confusion comes from.
What's actually being sold
A lot of what's marketed as "AI" right now is sophisticated workflow automation — a universal adapter sitting between systems that used to need custom integration code and a human in the loop to bridge them. That's genuinely valuable. It's also a fundamentally different value proposition from "AI replaces cognition," and it has a much lower compute ceiling. Workflow automation doesn't need frontier-scale inference running around the clock; it needs good enough, cheap, and reliable. If that's really where most near-term ROI lives, it argues for smaller, cheaper, more efficient models winning the volume game — not for data centers scaling at their current rate indefinitely.
The market is already telling on itself here. Worldwide IT spending is projected to hit $6.31 trillion in 2026, up more than 13 percent year over year, and Big Tech capital expenditure on AI has been reported north of $740 billion this year — a roughly 69 percent jump from 2025. Against that buildout, an Nvidia vice president told Axios flatly: "For my team, the cost of compute is far beyond the costs of the employees." Uber's CTO said the company burned through its entire 2026 AI coding budget in four months, after 84 percent of its engineers adopted an AI coding tool and roughly 70 percent of committed code started coming from AI — and Uber's own COO said publicly that token usage didn't correlate with the value of what actually shipped. Microsoft's internal numbers reportedly show something similar. Meanwhile, Block, Salesforce, Meta, Oracle, and Snap all made AI-attributed layoffs this year — companies cutting headcount on the premise that the machine was cheaper, then discovering the compute bill sometimes exceeds the payroll it replaced.
There's a name for the culture that produced this mismatch: "tokenmaxxing." Some AI providers subsidized inference costs early on, and usage-based pricing on top of artificially cheap access created incentives with no relationship to value delivered — Meta reportedly ran internal leaderboards where employees competed to burn through the most tokens. That's the same story client-server told about mainframe leasing costs, and cloud told about on-prem TCO: the sticker price quoted at adoption time excludes a category of hidden cost, and once someone actually measures total cost of ownership, the calculus flips.
What's new is the timing. Historically, the "we overbuilt this" reckoning took years to surface — the dot-com fiber glut, the on-prem TCO revolt, both played out well after the infrastructure was in the ground. This time, the bill is arriving before the buildout even finishes. That's either an early warning the industry can still act on, or the first tremor of a correction nobody's positioned for.
Quantum is not the rescue
The instinct, faced with a resource ceiling, is to reach for the next paradigm — and quantum computing gets cast as that paradigm more often than the physics supports. Quantum isn't a cheaper drop-in replacement for classical compute; it's a fundamentally different computational model, good at a narrow class of problems — certain optimization, simulation, factoring — that doesn't include the matrix multiplication underlying transformer inference. Error correction overhead currently means a single reliable logical qubit costs an enormous number of physical qubits, which makes the "cost-effective infrastructure" problem worse for quantum today, not better. The more likely outcome is quantum as a specialized co-processor for specific workloads — materials science, drug discovery, certain cryptographic and optimization problems — running alongside classical AI infrastructure, not replacing it. The more mundane resolution to the resource ceiling is the one that's actually worked before: efficiency gains in classical hardware, smaller models doing what large ones do now, and a market correction that prices usage closer to the value it delivers instead of a subsidized loss-leader rate.
The calculator problem, but for thinking
The infrastructure story has a self-correcting mechanism built in: money runs out, boards ask for ROI, capex gets rationalized. The cognitive story doesn't have an equivalent circuit breaker, and it's the part of this pattern most worth sitting with.
The calculator didn't kill math. It killed manual arithmetic as a bottleneck, and let people work at a higher level of abstraction — algebra, statistics, modeling — instead of grinding long division. But it's not a clean story of costless liberation. Mental-math fluency genuinely eroded; number sense, the ability to estimate and sanity-check an answer, has measurably degraded in students who reach for a device from an early age. What didn't degrade the same way was higher-order reasoning — arguably it rose, freed from the arithmetic it no longer had to carry. The calculator reallocated where the rigor lives. It didn't eliminate rigor; it moved the floor.
If AI does the same thing to writing, the honest question isn't "will people write worse sentences." It's whether writing has a load-bearing equivalent of arithmetic fluency — something foundational and mostly invisible until it's gone. I think it does, and it's a different kind of load-bearing than arithmetic ever was: the skill of forming a thought by writing it out. Arithmetic and reasoning are separable — nobody's mental model of the world degrades because they stopped doing long division by hand. Writing and reasoning may not be separable for most people. The friction of composition is frequently where the thinking happens, not something that occurs beforehand and then gets transcribed. Remove that friction the way a calculator removed long division, and you don't just risk losing polished prose — you risk quietly losing a piece of the reasoning process itself, for anyone whose habit of thought was built on the act of writing it down.
That's the real asymmetry, and it's worth stating plainly rather than resolving too neatly. It's also not a new fear. Socrates worried that writing itself would destroy memory and produce "the appearance of wisdom without the reality." Typewriters, word processors, spellcheck, and autocomplete each drew a version of the same complaint, and none of them ended up civilization-ending — though the skeptics weren't entirely wrong on narrower points either; spellcheck does correlate with worse unassisted spelling. The pattern holds across every one of these tools: some erosion at the floor is real and probably underway. What's undetermined is whether it stays contained to the mechanical layer, the way it did with arithmetic, or reaches into the cognitive layer the way it might with writing — and that depends less on the tool than on how people, and especially how children, are taught to use it before the underlying skill has had a chance to form.
Where this leaves the pendulum
Put together, the two cycles running through this moment don't resolve into a single forecast, and I don't think they should. The infrastructure cycle will almost certainly correct — not because AI stops being useful, but because the trillion-dollar bet on cognition-replacement is currently being funded by companies discovering the token bill exceeds the payroll it was meant to replace, and that math has a natural endpoint. What comes out the other side is likely a bifurcated architecture: concentrated, mainframe-scale compute for training frontier models, and a fragmenting, client-server-style ecosystem of cheap, efficient, often local inference for actually using them — closer to what workflow automation needs than what today's data-center growth rate assumes.
The cognitive cycle has no equivalent correction mechanism, because nobody sends an invoice when a generation quietly loses the habit of thinking by writing. That's the thread most worth pulling on next — not whether the data centers were overbuilt, but whether we notice the floor moving before it's gone.