The $20 chatbot subscription has not gone away, but it buys less than it used to. Firms that rely on cloud AI for real work are increasingly finding that the plan on the homepage is not the plan they actually need, and the gap between the two is where a local LLM starts to look attractive.
This is not simply a story about price rises. It is about what sits inside the cheap tier, who is funding the infrastructure behind it, and why running a model on hardware you own has stopped being a hobbyist pursuit.
The Entry Price Has Not Moved, but the Ceiling Has
ChatGPT Plus still costs $20 a month. What has shifted is everything above it. OpenAI introduced a $100 Pro tier on 9 April 2026, positioned below its existing $200 Pro plan. Anthropic's Claude Max is available at $100 and $200. Google's Gemini now offers a $100 tier alongside its $200 Ultra plan. A firm doing serious, all-day work is typically looking at five to ten times the entry price, not because the $20 tier disappeared, but because it no longer covers the work that matters.
The monthly headline price is only part of it. OpenAI sells additional credits once the included allowance runs out, covering tools such as Codex, ChatGPT Work, and Excel integration, none of which are included in the base plan. Google's AI Pro and Ultra tiers work the same way, with extra credits available for purchase and unused monthly credits lost rather than carried over.
The subscription still gets you in the door. The heavy lifting increasingly arrives as a second, separate bill as limits tighten and features move behind metered credits. That is a genuine squeeze, even if it stops short of scrapping the subscription altogether. You still have a chatbot, but you no longer have unlimited useful output at a fixed, round price. That is exactly why a local LLM starts to look like a line item worth budgeting for rather than a side project, and it is a different calculation from the one most small firms face when they first bring AI into their workflow without a dedicated development team.
Why a $20 Plan Cannot Cover the Infrastructure Bill
The other half of this story is capital expenditure. Alphabet's second-quarter 2026 results put full-year capital spending at $195 billion to $205 billion, with $44.9 billion spent in that quarter alone, most of it on AI infrastructure. Amazon, Meta, and Microsoft have each guided towards similarly large figures this year.
A $20 seat was never going to fund spending on that scale. Cheap consumer plans are currently being subsidised by investors, and investors expect returns eventually. When that pressure lands, prices rise, allowances shrink, and more of the genuinely useful work moves onto usage-based billing. Running a local LLM is one practical way to step off that treadmill, since the cost becomes a fixed piece of hardware rather than an open-ended subscription that can be repriced at any point.
What a Local LLM Setup Looks Like Today
The clearest recent example is Apple's Mac Studio update, announced on 25 August, which is aimed squarely at professional buyers rather than hobbyists. The new range runs on M5 Max and M5 Ultra chips. M5 Max reaches 128GB of unified memory with 614GB/s of bandwidth on its 40-core GPU configuration, while M5 Ultra scales to 512GB of unified memory and 1.2TB/s of bandwidth. Pre-orders are open, with machines shipping from 22 September and the 512GB configuration following in late October. Pricing in the United States starts at $2,499 for M5 Max and $5,499 for M5 Ultra.
Apple markets the machine on the basis that it can handle very large models locally, with full privacy and no token counting or creeping cloud costs. That positions a local LLM as a genuine office fixture rather than a lab experiment. The base Max configuration is a capable workstation, but it is the Ultra with 256GB or 512GB of memory that is aimed at running large models on-device. Firms needing more headroom can cluster several machines over Thunderbolt 5, though that route suits specialist deployments more than the average small or medium-sized business.
Other Local LLM Hardware and Their Limits
Apple is not the only option, and the alternatives come with their own trade-offs.
NVIDIA DGX Spark
NVIDIA's DGX Spark is a 1.2kg unit built on GB10 Grace Blackwell silicon, with 128GB of LPDDR5x memory and 273GB/s of bandwidth. NVIDIA quotes up to one petaFLOP of FP4 performance with sparsity, a theoretical figure rather than a direct comparison with Mac Studio throughput. The company states it can run inference on models up to roughly 200 billion parameters, fine-tune models up to 70 billion parameters, and pair two units for models up to 405 billion parameters. It draws power through a 240W supply and has been shipping since 15 October 2025; NVIDIA has not published a fixed retail price.
The trade-off is capacity. Spark carries a quarter of the Mac Studio Ultra's memory and around a fifth of its bandwidth, making it better suited to a dense inference appliance than a general-purpose 512GB workstation.
AMD Ryzen AI Max+ 395
AMD's answer is the Ryzen AI Max+ 395, also known as Strix Halo, aimed at laptops and mini-PCs rather than desktop workstations. It packs 16 Zen 5 cores and a Radeon 8060S GPU with 40 compute units, supports up to 128GB of 256-bit LPDDR5x-8000 memory with 256GB/s of bandwidth, and runs at a TDP of 45W to 120W. AMD says the chip can handle 70 billion parameter models on-device, with up to 112GB available to the GPU, and quotes up to 126 TOPS including 50 NPU TOPS.
Taken together, a local LLM setup now comes in three practical forms: a Mac Studio Ultra for firms that need half a terabyte of memory in a single machine, a DGX Spark for those wanting a compact NVIDIA appliance geared towards inference and fine-tuning, and a Strix Halo mini-PC or laptop for anyone who finds 70 billion parameters at 128GB sufficient. None of these are free, but each converts an open-ended monthly bill into a known, one-off capital cost.
Privacy Is a Stand-alone Reason to Go Local
Cost is the easier story to tell, but for lawyers, accountants, clinicians, and other regulated professionals, privacy is often the deciding factor. Using a public chatbot means prompts leave the building, and even where a vendor promises not to train on submitted data, the file itself still travels externally. Some clients will not accept that arrangement, and some regulators will not permit it either, which makes a local LLM a compliance decision as much as a financial one.
Keeping the model and the associated documents on hardware under direct control matters for legal privilege and for health records, and for any organisation unable to send client files to a consumer application hosted overseas. This also narrows the exposure firms have to a wider category of risk around AI systems being targeted or misused, which is a separate concern from cost but sits alongside it. Running locally does not remove the need for patching and access control, but it does remove the requirement to send privileged material to a shared third-party endpoint just to get a usable answer.
Mixing Local and Hosted Tools Sensibly
Adopting a local LLM does not mean abandoning hosted tools altogether. Most firms will end up running both: on-device models for drafting, searching internal files, and anything that should not leave the building, alongside a hosted assistant for public-facing tasks such as website chat, where the underlying knowledge is meant to be shared openly. That split looks more like a straightforward integration than a bespoke, from-scratch AI build.
A branded, hosted assistant trained on a firm's own content is a different job entirely from an air-gapped research machine, and the choice of partner for that public-facing layer still matters. It is also worth being realistic about what any of this hardware replaces: a faster drafting tool is not a substitute for professional judgement, and someone still needs to check the output regardless of where the model runs.
Not every task belongs on local hardware either. A public chatbot needs to be reachable around the clock, and the hosting, updates, and support for that layer will still sit with a partner for most small and medium-sized firms. A local LLM handles the private, internal side of the work; it does not replace the public one.
The practical question for most businesses is no longer whether a local LLM can do the job, but which parts of the workload should move on-premise and which should stay hosted. If you're a business owner and need help to decide how this hardware/cloud split should look, reach out to us and we will offer the advice and expertise you need to make informed decisions.