AI Economics

AI Pricing in 2026: Pro Limits, Local Qwen, and Cost per Job

Is AI getting cheaper, or are subscriptions getting worse? Alex Grankin's answer is that both can be true. He follows the September changes to ChatGPT Pro, compares a small local Qwen model with paid frontier models, and asks whether a dedicated Mac is now a sensible way to escape recurring fees. The most useful takeaway is not a single price chart: choose the least expensive setup that reliably completes your particular job at an acceptable speed.

The short answer

OpenAI confirms that the newly available $200 Pro plan has a lower included allowance for new and non-grandfathered subscribers; eligible existing users retain their prior allowance through 29 October 2026. Meanwhile, Qwen3.8-27B's open weights create a genuine local option for some jobs. Neither development means that every subscriber lost exactly half their usable work or that a local model replaces the best paid models across tasks.

Watch the Original Analysis

Credit: Alex Grankin's video, published 1 October 2026, and the supplied transcript. The comparison and creative-model judgments below are his observations; plan rules and model specifications are checked against the linked original sources. Alex also publishes at The AI Bridges.

What Changed in Paid AI Plans?

OpenAI's Pro tier guidance lists Pro 100, Pro 200, and Pro 500. Pro 200 reopened at $200 per month with a lower allowance for new subscriptions. Eligible earlier subscribers keep their previous allowance through 29 October 2026, then move to the revised one. Pro 500 is $500 per month and includes Astra Ultrafast. These are subscription allowances, not prepaid API tokens. OpenAI says work completed depends on the selected model, task, and settings rather than a universal number of messages.

At the API, OpenAI lists GPT-6 Astra at $10 per million input tokens and $50 per million output tokens for standard short-context text. The cheaper Sol and Luna models have different rates; longer contexts, tools, and processing modes can change a bill. Anthropic lists Fable 5.1 at $10 input and $50 output per million tokens, with lower cache-read pricing. Anthropic now says Fable 5.1 is available to Pro, Max, Team, and Enterprise users, so Alex's history of earlier Fable access should not be read as the current entitlement for every account.

The $14,000 Number Is Not an API Credit Balance

In June, SemiAnalysis bought plans and ran long-horizon coding tasks until weekly limits were exhausted. It priced the resulting usage at API list rates and estimated that the then-current $200 ChatGPT Pro and Claude Max plans could represent up to roughly $14,000 and $8,000 of API-equivalent retail value under that stress test. It did not report an API wallet that subscribers can spend, an average customer's usage, or either lab's marginal serving cost.

Alex describes the revised Pro allowance as roughly half the old allowance measured in API-price dollars. OpenAI's current plan page confirms a lower included allowance, but does not promise a fixed number of tokens or jobs. If lower-priced models improve, the same dollar-denominated allowance can still finish more routine work. A heavy Astra user and a lighter Sol user can therefore experience the change differently. Treat a historical $14,000 estimate as context, not today's personal plan balance.

Where Local Qwen Helps, and Where It Stalls

Qwen3.8-27B is a 27-billion-parameter, Apache-2.0 open-weight vision-language model. Running a suitable quantized build on your own machine can keep inference local after download, but "free model" does not mean free hardware, electricity, setup, maintenance, or review. Whether it fits and responds quickly depends on model file, quantization, memory, context length, and machine. Alex's LLM Sizer is a starting estimate, not a substitute for a timed test on your own computer.

In his Qwen-versus-Opus-4.6 creator test, Alex preferred Qwen on several short visual and coding examples. He also reports slower local iteration. In a later three-model face-off, he found the local model less convincing on complex game and 3D builds, while a weather app was closer. These are interesting samples, not a general quality ranking. Independent NIST/CAISI tests on a different open-weight model illustrate why public benchmark proximity can overstate performance on held-out tasks; they do not directly score Qwen3.8-27B.

Qwen3.8 vs Opus 4.6

Alex's earlier local-versus-frontier test.

Opus 5.5, GPT-6 Sol, and local Qwen

Alex's later creative-workflow comparison.

Compare Cost per Accepted Job, Not Just Tokens

API prices keep moving, but a cheaper token is not automatically a cheaper outcome. An agent may consume more reasoning tokens, retry more often, or need more human correction. Artificial Analysis measures the cost of completing defined benchmark tasks and finds large differences between models; its results are task-specific, not quotes for your workload. Run a small evaluation using your own examples:

MeasureWhat to recordWhy it matters
Accepted outputJobs approved without major correctionSeparates usable work from attractive demos.
Full spendPlan or API cost, local machine amortization, electricity, tools, and retriesStops a zero-token-price model looking free.
Human timeSetup, waiting, review, repair, and supportSlow local iteration can dominate the bill.
RiskData location, retention, access, and failure consequencesPrivacy and reliability affect the acceptable route.

Working formula: total model and infrastructure spend plus human labor, divided by accepted jobs. Compare a subscription, a usage-based API, a local model on hardware you already own, and a hybrid route. Send only the difficult or failed cases to an expensive model when policy and data permissions allow it.

Should You Buy a Mac for AI?

For most people, not just to cancel a subscription. Alex reaches a similar near-term conclusion. First test a local model on existing hardware and record its speed and acceptance rate on the work you actually do. Then compare a rented GPU, a dedicated machine, and a paid model using expected monthly utilization. A Mac can make sense for recurring private inference or media work that keeps it busy; it is a poor hedge if your important tasks still need a frontier model and the machine sits idle. Our M5 Ultra workflow test and cloud-versus-home hardware guide go deeper on that decision.

Alex speculates that subscription allowances and API prices may converge as models get more efficient. That is a forecast, not a published commitment about future plan value. His video also links subscription changes to possible public-market pressure; without a filing or company statement that establishes causation, treat that as his interpretation rather than a verified reason for a specific price change.

Turn the Pricing Shift Into a Business Idea

The opportunity is not a generic "cheap AI" service. It is a repeated customer task you can complete reliably with the right mix of human review, local processing, and paid models. This prompt keeps the buyer and quality threshold ahead of the hardware.

Business idea prompt

Find a job worth routing

Three offers, one measured seven-day test.

Ready to copy

Video Chapters

TimeTopic
00:00The pricing turning point
01:22Pro 200's revised allowance
01:54Subscription versus API-equivalent value
04:28How paid plans evolved
08:01Qwen3.8-27B versus Opus 4.6
12:03Reasons prices fall
13:38Price per token versus price per task
15:12Why frontier work remains expensive
20:00Subscriptions or pay-as-you-go?
20:50Should you buy a Mac?

Sources and Useful Links

Common questions

Did ChatGPT Pro lose half its included usage?
OpenAI says new or non-grandfathered Pro 200 subscriptions have a lower allowance at the same $200 price. Eligible existing subscribers keep their earlier allowance through 29 October 2026. The practical impact depends on model, task, and settings; it is not a fixed message or token count.
Does a $200 subscription include thousands of dollars in API credits?
No. SemiAnalysis estimated the API list-price equivalent of exhausting subscription limits with long coding tasks. That is not an API credit balance and is not the provider's actual cost to serve a user.
Can Qwen3.8-27B replace Claude or ChatGPT?
It can be useful for some private, routine, and well-scoped tasks on suitable hardware. Alex's small creative tests were mixed, and stronger frontier models still matter for complex, long-running work. Test quality, speed, and review effort on your own jobs.
Should I buy a Mac Studio to avoid AI subscription fees?
Not on subscription savings alone. Compare the hardware purchase, electricity, maintenance, utilization, model speed, human review, and privacy requirements against your present subscription and API use. Start on hardware you already own.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call