Local AI

Qwen 3.8 27B at Home: Speed, Cost, and the Real Trade-Off

Can a model you run at home change the economics of AI? Caleb Writes Code uses Alibaba's Qwen3.8-27B to make that case. His comparison is useful because it shows two real local configurations, not just a benchmark chart. It also illustrates why a cheap output-token calculation is only the beginning of the decision.

The short answer: Caleb reports about 41 tokens per second from a 3-bit quantized, roughly 13GB model on a dedicated GPU, versus four tokens per second for full-precision weights on his DGX Spark. At his stated 250W GPU draw and $0.21/kWh, the first setup uses about $0.36 in GPU electricity per million generated tokens. That is not its total operating cost, and the two setups do not run the same precision.

Watch Caleb's Qwen3.8-27B Breakdown

Credit: Caleb Writes Code's original video, published 26 August 2026, and the transcript supplied with this post. Caleb shares more work on X, LinkedIn, and TikTok. His video includes a paid DevRev Computer segment; it is disclosed separately below.

Why This 27B Model Gets Attention

The official model card identifies Qwen3.8-27B as a dense, open-weight text-and-vision model under the Apache 2.0 license, with a native 262,144-token context window. The Qwen repository announced the 27B release on 14 August 2026. The ability to download weights and choose a local runtime gives developers more control over deployment than an API-only model. It does not make every application private, inexpensive, or production-ready by default.

Caleb uses a then-current Code Arena placement to illustrate how competitive a smaller model can be. Leaderboard ranks move, and the Qwen benchmark table is vendor-reported under particular harnesses. Neither is a substitute for testing the workflows you actually need. For a coding example with observed corrections and completed artifacts, see Gary Explains' separate Qwen3.8-27B test.

Two Machines, Two Different Model Files

  • Dedicated GPU at home: About 13GB, 3-bit quantized weights. Caleb reports roughly 41 output tokens per second. The smaller file fits in fast GPU memory, with a precision trade-off to test.
  • DGX Spark with unified memory: About 55GB, full-precision weights. Caleb reports roughly four output tokens per second. Larger weights and a different memory path make this a different workload.

Caleb points to memory bandwidth as a major explanation, while acknowledging that his comparison is not apples to apples. Both the hardware and precision changed. The result does not prove that one system is always faster or that 3-bit output is equally good. Prompt length, context size, inference settings, and runtime overhead could also change the outcome. NVIDIA's DGX Spark specifications describe the platform; Caleb's speeds remain creator measurements, not NVIDIA or Qwen performance guarantees.

What Does a Million Local Output Tokens Really Cost?

Caleb's arithmetic is reproducible as a GPU electricity estimate: one million output tokens at 41 tokens per second takes about 6.8 hours. At 250 watts and $0.21 per kilowatt-hour, that is roughly $0.36, close to his quoted $0.37. He contrasts it with a hosted listing at about $3 per million output tokens and describes an approximately eightfold gap.

That gap is real within the narrow comparison shown, but it is not a like-for-like total-cost comparison. The local run uses a 3-bit model; the hosted price he cites is for a different precision. A business decision should also count:

  • Hardware: GPU purchase or depreciation, other computer components, storage, repairs, and idle capacity.
  • Energy: the rest of the PC, cooling, and energy used before output generation, not just the GPU's stated draw.
  • Work: setup, updates, security, monitoring, failed runs, and human review.
  • Usage: input tokens, long prompts, context memory, concurrency, and whether you actually generate enough output to spread fixed costs.

For a low-volume user who would need to buy a GPU, a subscription or hosted API may still be cheaper. For a team with suitable hardware already running and a repeatable high-volume task, local inference may be attractive. Compare cost per accepted job and quality, not just cost per million output tokens.

Quantization: Fit First, Then Test Quality

Quantization stores model weights at lower precision so the file uses less memory. That can put a 27B model within reach of hardware that cannot hold its full-precision weights. Unsloth's Qwen3.8-27B guide lists roughly 13-16GB total memory for its 3-bit variants, 17-19GB for recommended 4-bit builds, and about 56GB for BF16. Those figures describe its published variants, not a guarantee for Caleb's exact file or your full runtime configuration.

Caleb shows a distribution-distance chart for quantizations. It is a useful warning that smaller files can change token probabilities, but a single divergence number cannot tell you whether your documents, code, or tool calls still pass. Keep a small independent test set, run the same tasks on the quantizations you can fit, and record quality, elapsed time, memory use, and intervention count.

The Subscription Also Pays for a Workflow

Caleb makes a second point that gets lost in model-price comparisons: coding and work subscriptions bundle the harness around the model. Histories, files, skills, connectors, team context, and scheduled work can be harder to replace than inference itself. A local model may be one component in a hybrid workflow rather than a reason to cancel every hosted tool. Start with a bounded task, check what data the runtime and integrations transmit, and preserve a human approval step for externally visible results.

Sponsored segment: From 04:15, Caleb presents DevRev Computer as a shared AI teammate with memory and integrations. DevRev sponsored the original video. Its product claims and the illustrated pitch-deck workflow are a sponsor demonstration, not a Qwen benchmark or an independent test by this site.

Turn This Into a Business Idea

The most promising offer may be a privacy-conscious workflow you know how to deliver, not a promise of cheap tokens. This prompt starts with your skills and a buyer's repeated job, then tests whether local processing changes the outcome enough to matter.

Business idea prompt

Find a local-first offer worth validating

Turn model economics into a buyer-specific experiment, with full costs and human review.

Ready to copy

Video Chapters

TimeTopic
00:00Why consumer-GPU models matter
00:29Model-release landscape
01:12Caleb's economics thesis
03:14Subscriptions and switching costs
04:15Sponsored DevRev Computer segment
05:45Community quantizations
06:29Dedicated GPU versus DGX Spark
07:56Precision and memory trade-off
08:49Electricity calculation and conclusion

Sources and Useful Links

Common questions

Is Qwen3.8-27B free to run locally?
The official weights are available under Apache 2.0, but running them still uses hardware, electricity, storage, setup time, and maintenance. A suitable quantized version may run on a consumer PC; the full-precision build needs much more memory.
Did Caleb prove that local Qwen is eight times cheaper than cloud inference?
No. His roughly $0.36 to $0.37 per million output tokens estimates GPU electricity only and compares a 3-bit local build with a hosted full-precision listing around $3 per million output tokens. It excludes hardware depreciation, other system power, input processing, and operational costs.
Why was the DGX Spark slower in the video?
Caleb reports around four tokens per second on the Spark with full-precision weights and around 41 on a dedicated GPU with a much smaller 3-bit model. The comparison changes both memory path and quantization, so it does not isolate one cause or rank the machines for every workload.
Does a smaller quantized file mean the same answer quality?
No. Quantization reduces memory use and can change model outputs. Benchmark a candidate quantization on your own tasks, especially errors, long context, tool use, and review time. A distribution-distance chart is not a substitute for task-level evaluation.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call