Is AI getting cheaper, or are subscriptions getting worse? Alex Grankin's answer is that both can be true. He follows the September changes to ChatGPT Pro, compares a small local Qwen model with paid frontier models, and asks whether a dedicated Mac is now a sensible way to escape recurring fees. The most useful takeaway is not a single price chart: choose the least expensive setup that reliably completes your particular job at an acceptable speed.
OpenAI confirms that the newly available $200 Pro plan has a lower included allowance for new and non-grandfathered subscribers; eligible existing users retain their prior allowance through 29 October 2026. Meanwhile, Qwen3.8-27B's open weights create a genuine local option for some jobs. Neither development means that every subscriber lost exactly half their usable work or that a local model replaces the best paid models across tasks.
Watch the Original Analysis
Credit: Alex Grankin's video, published 1 October 2026, and the supplied transcript. The comparison and creative-model judgments below are his observations; plan rules and model specifications are checked against the linked original sources. Alex also publishes at The AI Bridges.
What Changed in Paid AI Plans?
OpenAI's Pro tier guidance lists Pro 100, Pro 200, and Pro 500. Pro 200 reopened at $200 per month with a lower allowance for new subscriptions. Eligible earlier subscribers keep their previous allowance through 29 October 2026, then move to the revised one. Pro 500 is $500 per month and includes Astra Ultrafast. These are subscription allowances, not prepaid API tokens. OpenAI says work completed depends on the selected model, task, and settings rather than a universal number of messages.
At the API, OpenAI lists GPT-6 Astra at $10 per million input tokens and $50 per million output tokens for standard short-context text. The cheaper Sol and Luna models have different rates; longer contexts, tools, and processing modes can change a bill. Anthropic lists Fable 5.1 at $10 input and $50 output per million tokens, with lower cache-read pricing. Anthropic now says Fable 5.1 is available to Pro, Max, Team, and Enterprise users, so Alex's history of earlier Fable access should not be read as the current entitlement for every account.
The $14,000 Number Is Not an API Credit Balance
In June, SemiAnalysis bought plans and ran long-horizon coding tasks until weekly limits were exhausted. It priced the resulting usage at API list rates and estimated that the then-current $200 ChatGPT Pro and Claude Max plans could represent up to roughly $14,000 and $8,000 of API-equivalent retail value under that stress test. It did not report an API wallet that subscribers can spend, an average customer's usage, or either lab's marginal serving cost.
Alex describes the revised Pro allowance as roughly half the old allowance measured in API-price dollars. OpenAI's current plan page confirms a lower included allowance, but does not promise a fixed number of tokens or jobs. If lower-priced models improve, the same dollar-denominated allowance can still finish more routine work. A heavy Astra user and a lighter Sol user can therefore experience the change differently. Treat a historical $14,000 estimate as context, not today's personal plan balance.
Where Local Qwen Helps, and Where It Stalls
Qwen3.8-27B is a 27-billion-parameter, Apache-2.0 open-weight vision-language model. Running a suitable quantized build on your own machine can keep inference local after download, but "free model" does not mean free hardware, electricity, setup, maintenance, or review. Whether it fits and responds quickly depends on model file, quantization, memory, context length, and machine. Alex's LLM Sizer is a starting estimate, not a substitute for a timed test on your own computer.
In his Qwen-versus-Opus-4.6 creator test, Alex preferred Qwen on several short visual and coding examples. He also reports slower local iteration. In a later three-model face-off, he found the local model less convincing on complex game and 3D builds, while a weather app was closer. These are interesting samples, not a general quality ranking. Independent NIST/CAISI tests on a different open-weight model illustrate why public benchmark proximity can overstate performance on held-out tasks; they do not directly score Qwen3.8-27B.
Qwen3.8 vs Opus 4.6
Alex's earlier local-versus-frontier test.
Opus 5.5, GPT-6 Sol, and local Qwen
Alex's later creative-workflow comparison.
Compare Cost per Accepted Job, Not Just Tokens
API prices keep moving, but a cheaper token is not automatically a cheaper outcome. An agent may consume more reasoning tokens, retry more often, or need more human correction. Artificial Analysis measures the cost of completing defined benchmark tasks and finds large differences between models; its results are task-specific, not quotes for your workload. Run a small evaluation using your own examples:
| Measure | What to record | Why it matters |
|---|---|---|
| Accepted output | Jobs approved without major correction | Separates usable work from attractive demos. |
| Full spend | Plan or API cost, local machine amortization, electricity, tools, and retries | Stops a zero-token-price model looking free. |
| Human time | Setup, waiting, review, repair, and support | Slow local iteration can dominate the bill. |
| Risk | Data location, retention, access, and failure consequences | Privacy and reliability affect the acceptable route. |
Working formula: total model and infrastructure spend plus human labor, divided by accepted jobs. Compare a subscription, a usage-based API, a local model on hardware you already own, and a hybrid route. Send only the difficult or failed cases to an expensive model when policy and data permissions allow it.
Should You Buy a Mac for AI?
For most people, not just to cancel a subscription. Alex reaches a similar near-term conclusion. First test a local model on existing hardware and record its speed and acceptance rate on the work you actually do. Then compare a rented GPU, a dedicated machine, and a paid model using expected monthly utilization. A Mac can make sense for recurring private inference or media work that keeps it busy; it is a poor hedge if your important tasks still need a frontier model and the machine sits idle. Our M5 Ultra workflow test and cloud-versus-home hardware guide go deeper on that decision.
Alex speculates that subscription allowances and API prices may converge as models get more efficient. That is a forecast, not a published commitment about future plan value. His video also links subscription changes to possible public-market pressure; without a filing or company statement that establishes causation, treat that as his interpretation rather than a verified reason for a specific price change.
Turn the Pricing Shift Into a Business Idea
The opportunity is not a generic "cheap AI" service. It is a repeated customer task you can complete reliably with the right mix of human review, local processing, and paid models. This prompt keeps the buyer and quality threshold ahead of the hardware.
Find a job worth routing
Three offers, one measured seven-day test.
Video Chapters
| Time | Topic |
|---|---|
| 00:00 | The pricing turning point |
| 01:22 | Pro 200's revised allowance |
| 01:54 | Subscription versus API-equivalent value |
| 04:28 | How paid plans evolved |
| 08:01 | Qwen3.8-27B versus Opus 4.6 |
| 12:03 | Reasons prices fall |
| 13:38 | Price per token versus price per task |
| 15:12 | Why frontier work remains expensive |
| 20:00 | Subscriptions or pay-as-you-go? |
| 20:50 | Should you buy a Mac? |
Sources and Useful Links
- Creator analysis and tests: Alex Grankin's main video, Qwen versus Opus 4.6, and three-model face-off. His LLM Sizer, newsletter, and X profile are linked for follow-up.
- Subscription and API terms: OpenAI Pro tiers, Astra usage guidance, Astra API model page, and Anthropic's Fable 5.1 launch. Plans and rates can change; check them before purchase.
- Independent and original research: SemiAnalysis's June plan-exhaustion test, Artificial Analysis on cost per task, and NIST/CAISI's held-out evaluation.
- Local model: Qwen3.8-27B model card for weights, license, context, and deployment notes. Its hardware needs depend on the version you run.