AI Doesn’t Compute an Answer. It Guesses One, and We’re Paying for the Guess.
First published on substack Sept 4th 2026
I run Powerverse and like every other technology company we are gradually becoming reliant on Frontier AIs like Anthropic and OpenAI. Getting this bet wrong would destabilise the business. What follows are my conclusions on that bet when I went looking past their marketing.
Three things I’m now confident about
Don’t stop hiring junior staff to replace them with AI. It isn’t accurate enough to make that trade safely.
Do worry about the pricing. What providers charge doesn’t look like it carries the full cost of running these systems.
It looks quite speculative, in the same way as the money that built up before the dot-com correction.
The pitch sounds like magic, and I’ve heard that trick before
The current pitch for frontier AI reads like magic: describe the outcome of unfettered intelligence on tap, skip the mechanism of how we get there, and count on nobody asking how it actually works.
I’ve watched that trick from the inside before, as a competitor. I ran CloudSense an enterprise SAAS business in a then hot market category, and we used to sell against a San Francisco start up that raised half a billion dollars, taught its teams to lie about what the product could and would actually do, eventually leading to that company’s failure and fire sale. That wasn’t the only time I saw the trick. I think this stuff seems to happen all the time. For example I saw the same thing played out at a much bigger scale NYSE listed SaaS giant that launched a whole new “Cloud” suite at its annual conference to a hundred thousand people and live streamed worldwide, just before that product’s brief existence and inevitable failure. They were selling a dream built on a financially unsound technical nightmare. So excuse my cynicism when things are starting to seem too good to be true in the land of AI.
I’m not the one writing the code any more, though the engineering is never far from what I do as startup CEO. The last time I wrote it myself was around the Millennium bug, which turned out to be hyped up in a way that caught everyone by surprise - the same shape as the sales pitch above, just twenty five years earlier. So it is the same cynicism that led me to draw the 1990s telecoms parallel in a recent article I wrote.
It’s not computing an answer. It’s guessing one
Most of us would say with some certainty that a large language model doesn’t calculate the answer to our questions straight away, even now after years of trying. My moderate tech ability allows me to understand that this is because it samples answers from a probability distribution over plausible next steps, an approximation, not a lookup. That’s true of the cheapest chatbot and the priciest frontier model alike, and what I’ve been learning more recently is that this is the starting fact behind everything else here: why the economics look the way they do, why we don’t take its answers as guaranteed when the work has to be accurate (unlike excel), and why I don’t buy the version of the future where one of these systems simply takes over. A system built to approximate, it turns out isn’t a system built to be certain, however large it gets. This is the same distinction Cal Newport draws describing language models on the Better Offline podcast, 26 August 2026. This was the silver bullet opportunity that my CTO and others have been wrestling with in the quest for productivity and it turns out a skilled human in the loop is mission critical.
For a great many of the problems being handed to these LLM systems, a deterministic method already exists that reaches a guaranteed answer faster and for less: forecasting, optimisation, ordinary arithmetic. Using a probabilistic model for a job that already had a deterministic answer, then spending to compensate for that choice, is the misplaced decision underneath everything below.
Making that approximation reliable enough to sell it to us customers, costs money twice over. Once in the harness around it - retrieval, tool calls, an agent checking its own work - which runs on every query and never gets cheaper just because it ran before. And once in training: curated data, human feedback, compute spent adjusting billions of parameters, meant to be paid up front and spread across every future query the way a factory’s construction cost spreads across everything it later makes. Agentic work already consumes five to thirty times the tokens of a simple chat exchange for exactly this reason, hundreds of times more on demanding coding tasks. (Spheron Network, citing Stanford research on agentic token consumption.)
What that decision costs
None of this was easy to untangle. Frankly I didn’t understand it, but with much re-reading and re-listening to the excellent analysts that are not heralding the dawn of computer consciousness here’s what I found out..
Training a new model - a necessity for every new use case (they don’t train themselves) hasn’t got cheaper, because narrowing LLM approximation is exactly what most of the money is spent trying to do. Dario Amodei described frontier runs going from “on the order of $100 million” to “closer to $1 billion” to “$5 or $10 billion” inside about two years, and Epoch AI puts the underlying growth rate at 2.4 times a year. (The Decoder, on Amodei’s April 2024 interview; Epoch AI.)
On the inference side; which is the day to day LLM problem solving process, the token price really has been falling fast, around 50 times a year since early 2024, (based on Stanford’s AI Index data) but that is not a forward looking statement on cost and it does not help the ‘cost of approximation’ problem underneath it. Because nobody buys tokens. They buy a harness built to compensate for that accuracy gap, and Gartner forecasts the true cost per agentic workflow will rise more than fivefold through to 2028. (Gartner, 17 August 2026.) The “compute is getting cheaper” story for inference rests entirely on the token cost and doesn’t describe the rising cost of LLM procedural complexity at run time, that hits the customer’s bill.
The costs of all this are not concealed once you look, it’s just split across two sets of books.
OpenAI paid Microsoft $17.2bn for compute in 2025; roughly $10.6bn went down as R&D and $6bn as the direct cost of serving customers, which is how it reports a 43% gross margin and a $20.9bn operating loss on $13.07bn of revenue in the same year. (Ed Zitron, “Exclusive: OpenAI Losses Increased Nearly 8X in 2025,” 16 June 2026, independently verified by the Financial Times.)
On Microsoft’s own books the same spend shows up as depreciation: cloud gross margin fell to 66%, free cash flow fell 23% in the fourth quarter, and Alphabet’s and Amazon’s have already gone negative. Roughly two-thirds of that capital spending is GPUs, which depreciate on a three-to-six-year clock against a building’s 25 years. (Microsoft FY2026 Q4 earnings, 29 July 2026; Investing.com.) That means those providers have three to six years to generate the vast revenue to pay for those GPUs.
Who is actually footing the bill?
A $200-a-month subscription doesn’t come close to covering the amount of compute the heaviest users can consume. Something I found out recently when I used Fable from Anthropic and burned through a week’s worth of tokens in a day (and then started researching this article!)
One study by AI/Semiconductor industry research firm SemiAnalysis in June bought every consumer tier OpenAI and Anthropic sell and ran each flat out until it hit its usage limit. The top ChatGPT plan could consume roughly $14,000 of compute at published rates, the equivalent Claude plan around $8,000, against a developer rule of thumb closer to $2,000. (SemiAnalysis, June 2026; summarised by TechSpot.)
Even adjusted for list-price margin, the real figure is nearer $3,500. When Anthropic tried metering that gap directly in May 2026, it paused the change on the day it was due to take effect. (Fochis; InfoWorld; DevOps.com.)
Follow where the money that’s supposed to deliver the AI bet is actually coming from and a circular pattern appears.
Microsoft puts money into OpenAI. OpenAI then spends almost all of it straight back on renting Microsoft’s cloud computers: $17.2bn went from OpenAI to Microsoft in 2025, while only $303m came back the other way. Microsoft counts what OpenAI pays it as revenue, and uses that revenue to justify spending even more on AI infrastructure. (Ed Zitron, same reporting as above.)
This is where my dot-com memory earns its place in this article as an ‘aide memoire’ rather than an anecdote.
As the Bank of England says - AI-linked companies now make up roughly half of the S&P 500, up from about a quarter in 2022. Their own own stress scenario models a 45% fall in US equities over six quarters if that concentration takes a shock. (Bank of England, Financial Stability Report, July 2026.)
That would be the telecoms build-out story all over again, just with better marketing this time around.
How Powerverse is starting to hedge these risks
If a big AI vendor fails, gets restructured, acquired or wound down the way dot-com casualties were, we need a hedge against it. The same goes if the cost curve simply keeps climbing instead of levelling off. Either one is enough to leave a business like mine stranded if it has built its core operations exclusively on those frontier models. The root problem can’t be glossed over: a system that guesses doesn’t become a system that knows just because a trillion dollars makes it a better guess. So the mitigation is multilayered: our models choices for the outcomes that have to be guaranteed - forecasting, optimisation, control cannot be LLM ones. More and more, open-source and cheaper models must cover the jobs where a language model is the right tool but a frontier price tag isn’t. Frontier models get used only where they earn their place. Cal Newport sums this position up as “multiple components, some symbolic, some neural,” rather than one vendor’s model wired into everything we do. (Quoted from the same Better Offlineepisode.) This is the sort of hedge we are moving towards. Not to mention hiring. We won’t stop hiring smart, entry level colleagues any time soon.
Then if these vendor’s pricing changes, or the funding runs out, or “a lot of the more impressive stuff” turns out to come from smaller, “super tuned” systems instead, our exposure to any one of those outcomes is limited by design. With that mitigation, we can stay on track whatever happens in Frontier AI.