9 min read

Can You Cancel Your AI Subscription Yet? I Did the Math.

A viral video says free AI models you run yourself now give you 87% of the quality at 1% of the cost. The quality part is true. The savings part falls apart the moment you look at your electricity bill.

A paper receipt itemising 30 cents of electricity to run an AI model at home against 13 cents to rent the same model per million tokens, a 2.3x difference

A video went round my feed this week telling everyone to cancel their AI subscriptions. Free models you run on your own computer now give you "87% of the fidelity at 1% of the cost," it says, so the big AI companies are "totally and completely cooked."

If that is true, it is money back in your pocket every month. So I spent a day checking it.

The quality claim is basically right, and closer than I expected. The savings claim falls apart the moment you look at your electricity bill.

The free models really did catch up

Quick bit of vocabulary first, because the video skips it. A free or "open" model means the company published the actual file that makes it work, so anyone can download it and run it. ChatGPT and Claude are the other kind. They stay on the company's servers and you rent access by the month or by the word.

On the main independent scoreboard, the Artificial Analysis Intelligence Index (opens in a new tab), the best free models score 60 in early September 2026. The best paid model scores 66. That is 91%, so the video actually undersold it.

Two years ago the free ones were about a year behind. In July 2026 the UK government's AI Security Institute (opens in a new tab) tested the best free models on security work and found them matching paid models from four to seven months earlier, down from a six to ten month gap through 2025.

SemiAnalysis, who follow this for a living, published an analysis on August 21, 2026 (opens in a new tab) showing the catch-up time roughly halving with each new generation. That is remarkable and I do not want to talk anyone out of being impressed by it.

Here is the first thing the video leaves out. Free does not mean it fits on your laptop. The models scoring 60 are gigantic, somewhere between 744 billion and 2.8 trillion settings each, and they need a rack of specialist graphics cards.

You can rent them cheaply from a hosting company. You cannot run them at your desk.

The best free model that does fit on a good home computer scores about 52, which is roughly the level of the cheap tier of a paid service. Still useful. Not the frontier.

The second thing missing is that the gap is lopsided. Ask for an email, a summary or a chunk of ordinary code and you would struggle to tell the two apart. Hand over a job with fifty steps that all have to go right and they separate badly.

How long a model can work unsupervised before it goes off the rails is the thing that separates them, and it is also the hardest thing to measure. METR, the lab that times exactly that (opens in a new tab), now says its own tests stop being reliable past about 16 hours. The paid models are the ones pushing against that ceiling.

The free model is level with the paid one on a five-minute task and nowhere near it on a five-hour one. A single average score hides that completely.

That last stretch of reliability is where the money lives, because it is the difference between a demo that impresses your mate and something a paying customer touches. I have shipped both. The demo is the easy one.

Is it really 1% of the cost? Only if you pick the right two things

AI is sold by the million tokens, and a token is about three quarters of a word. So a million tokens is roughly 750,000 words going in and coming back out, which is about ten novels.

Here is what a million tokens costs across the market today.

Cost per million tokens, mixed at three words in for every one word out, from the official pricing pages of OpenAI (opens in a new tab), Anthropic (opens in a new tab), DeepSeek (opens in a new tab) and DeepInfra (opens in a new tab), checked September 3, 2026.
Model Type Cost
Claude Fable 5.1 Paid, top tier $20.00
Claude Opus 5 Paid, top tier $10.00
Kimi K3, best free model Free, rented $6.00
Claude Sonnet 5 Paid, mid tier $4.00
GPT-5.6 Luna Paid, budget tier $0.45
DeepSeek V4 Flash Free, rented $0.105

So the claim survives, barely. DeepSeek Flash really is about 1% of the price of a top-tier Claude, and for plenty of everyday work Flash is fine.

It is also the cheapest free model held up against the most expensive paid one. Line up the models that actually compete and the best free model costs 60% of the top Claude, and half again more than the mid-tier Claude most people should be using anyway. OpenAI's budget option is four times the price of DeepSeek Flash, not a hundred times.

Where "1% of the cost" is genuinely true is over time. The price of any given level of AI quality falls roughly ten times a year (opens in a new tab). Whatever felt like magic two years ago now costs pennies.

The top of the market never gets cheap though, because the top keeps moving. People pay for whatever the best is, and the best keeps changing.

Running it on your own machine costs more than renting it

This is the part everyone repeats and nobody checks.

Someone ran the careful test in August (opens in a new tab), on the best machine for the job, a $6,000 Mac Studio. The electricity alone came to about 30 cents per million tokens. Renting the exact same model from DeepSeek costs 13 cents.

Read that again, because it is backwards from what everyone assumes. Running the free model yourself cost more than twice as much as renting it from a company that still has to make a profit on you. And that is before a penny of the $6,000 computer.

The reason is boring. A data centre runs the same expensive hardware flat out, around the clock, shared between hundreds of customers. Your machine sits idle between your questions, and you paid for all of it.

When buying the machine does pay off

Small model, simple repetitive job, hardware you keep busy all day. One writer swapped a $1,500-a-month bill for a $2,500 Mac (opens in a new tab) sorting 45,000 items a day, and paid the machine off in three months. If that is your workload, buy the Mac. If you ask a handful of hard questions a day, rent.

There are good reasons to run your own model. Your data never leaves the building, it works with the internet down, and nobody can change the terms on you overnight, which is the argument I made about what owning your stack actually buys you. Saving money is not on that list.

Local AI already won, and you are already using it

Here is the funny part. The video is right that small models running on your own device are everywhere. They are just invisible.

Your phone keyboard's suggestions, your photo search, page summaries in Chrome, Apple's proofread button, the little assistant in Windows Settings. That last one runs on a model 8,000 times smaller (opens in a new tab) than the free ones topping the scoreboard, and it works fine, because changing a setting is not a hard problem.

Then look at where those same companies send the hard work. Apple's own design (opens in a new tab) hands anything genuinely complicated to a big model in a data centre, and Siri's brain is Google's Gemini. That is the shape of local AI in 2026. Small and everywhere for easy jobs, cloud for anything that has to think.

One number for scale. Google alone handles 3.2 quadrillion tokens a month. The year-long study everyone quotes to prove free models are winning covered 100 trillion tokens in total, which is about a day of Google.

The companies everyone said were finished had their best year

This is where the video stops being arguable and starts being wrong.

Company-reported figures, none of them audited yet. From TechCrunch (opens in a new tab), Sacra (opens in a new tab) and GeekWire (opens in a new tab).
Company Late 2025 Mid 2026
Anthropic $9B a year $65B a year
OpenAI $20B a year About $40B a year
Microsoft AI Not disclosed $37B a year, up 123%

Microsoft says it cannot build servers fast enough to meet demand. That is not the problem you have when customers are leaving.

Business spending tells the same story. Menlo Ventures found companies spent $37B on AI in 2025 (opens in a new tab), more than three times the year before, and the slice going to free models fell from 19% to 11%.

So why do you keep seeing headlines saying free models are taking over? Because they measure different things.

Hobbyists and developers push enormous volumes of text through whatever is cheapest, so free wins on word count. Businesses pay for the thing that does not fall over on a Tuesday, so paid wins on money. Money is what decides who survives.

The real risk, and the bit the video has backwards

There is a serious case that these companies are in trouble. It has nothing to do with you cancelling your subscription.

OpenAI burned $3.7B in three months (opens in a new tab) and has committed to something like $1.4 trillion of computing power over eight years. If that bet never pays back, the story is a company that overspent, not a customer revolt.

And the bit that is genuinely backwards. Free models are followers. Every leap they have made was aimed at a target a paid lab set first, and part of how they catch up is training on the paid models' own answers. Nobody has shown the free labs would push forward on their own if the expensive ones stopped.

"Free and open" has also quietly come to mean "Chinese." The models at the top of that scoreboard come from Alibaba, Moonshot, Zhipu and DeepSeek. Meta stopped publishing its best work. There is a bill in Congress (opens in a new tab) to ban Chinese models from federal agencies, and several US states and departments have already banned DeepSeek on work devices.

If you are about to build your business on those weights, that is worth knowing before you start, not after. I went into the politics of it when the benchmarks first flipped in June.

Free AI models vs paid subscriptions: FAQ

Is running AI on my own computer cheaper than paying a monthly subscription?

Almost never. Someone measured a top free model on a $6,000 Mac Studio in August 2026 and the electricity alone came to about 30 cents per million words of traffic, against 13 cents to rent the identical model from a hosting company. You save money only on small, repetitive, high-volume jobs on a machine you keep busy all day.

Are free AI models as good as ChatGPT and Claude now?

On short everyday tasks, close enough that most people would not notice. The best free models score about 60 on the main independent scoreboard against 66 for the best paid one. The gap opens up on long, complicated work with many steps, and the free models scoring 60 are far too large to run on a normal computer.

What computer do I need to run a good AI model at home?

Around $2,000 buys a mini PC that runs mid-sized models at reading speed. About $6,000 buys a Mac Studio that runs one of the big free models slowly. The models that genuinely rival the paid services want $10,000 to $20,000 of hardware and still have to be compressed to fit.

Are businesses actually switching from paid AI to free models?

The money says no. Menlo Ventures found the share of company AI budgets going to free and open models fell from 19% to 11% during 2025, even as total spending more than tripled. Free models do win on raw volume among hobbyists and price-sensitive developers, which is why you see headlines pointing both ways.

What I would actually do

Mix them, because that is what the arithmetic says. Cheap free model for the boring high-volume work, a mid-tier paid model for most things, the expensive one for the handful of jobs that have to be right. People who track their own spend land near 70% cheap, 25% mid, 5% premium, and pay about 15% of an all-premium bill.

Whatever you do, do not buy a $6,000 computer to get out of a $60 subscription. That is the one move in this whole story that reliably loses money.

Here is what would change my mind, and it is a specific machine. The M5 Ultra Mac Studio lands in October with 512GB of memory. If that runs one of the top free models at a speed you would actually put up with, the sums above stop being so one-sided and I will write the follow-up saying so.

I have one on my watchlist. Ask me again in November.

Ready to Turn These Insights Into Results?

Don't let technical debt, slow load times, or rigid templates bottleneck your business growth. Get a robust, custom technical architecture engineered specifically for your brand.