AI Model Reviews

Qwen 3.8 27B Locally: Gary Explains' Coding Test

Qwen3.8-27B is a 27-billion-parameter open-weight model that Gary Explains tested on his own PC. In a six-minute video, he moves past short reasoning questions and uses OpenCode to make a game, implement a binary protocol, and build an interpreter in C. The interesting result is not that a model wrote code. It is that he could keep the model on a local machine and still finish work that had stopped smaller local models in his earlier attempts.

The short answer: Gary reports about 96 generated tokens per second from a 4-bit quantized Qwen3.8-27B on his 32GB RTX 5090. His game, protocol implementation, and example interpreter ran, with human guidance along the way. One protocol specification example was wrong. This is strong evidence that the model was useful for his workflow, but not a controlled demonstration that it is the new leader for every local coder.

Watch Gary's Qwen3.8-27B Test

Credit: Gary Explains' original video, published 18 August 2026. The observations below come from the supplied transcript and his demonstration. Model details are checked against the official Qwen model card as of 1 October 2026. Gary is also on X and GitHub.

What Actually Ran on His PC?

Qwen's release is a dense 27B model with text and vision support, a native 262,144-token context window, and an Apache 2.0 license. Those are model properties, not a promise that a consumer GPU can use the entire context at a useful speed. The official project repository documents local serving options and points to quantized variants.

Gary estimates that the 16-bit weights are roughly a 55GB download, which does not fit wholly in the RTX 5090's 32GB of VRAM. He instead loaded a 4-bit quantization and measured about 96 tokens per second. That figure is his one-machine observation, not an official Qwen speed rating or an end-to-end coding time.

32GB is his tested setup, not a universal minimum. For example, a later Unsloth GGUF catalog lists a 4-bit variant around 17.5GB. The model file still needs headroom for runtime state, the KV cache, context, and other processes. A smaller card may require a different quantization, shorter context, or CPU offload, with different speed and quality. Do not infer Gary's 96-token result from file size alone.

What the Creator Tested

TaskObserved resultWhat it tells us
Short reasoning questionsAnswered the Alice and two-egg-timer questions.A warm-up, not a demanding coding evaluation.
Single-file Space Invaders gameProduced an index.html with inline JavaScript that Gary opened and played.It could deliver a small runnable artifact without external libraries in this test.
Type-length-value protocolDrafted a specification, but one or two worked examples had incorrect numbers.The design looked useful, yet its examples still needed verification.
Protocol library in OpenCodeImplemented the spec with unit and end-to-end tests after Gary supplied a plan and follow-up guidance.The model completed a multi-step coding job with an operator in the loop.
New Scrippy interpreter in CInferred behavior from example code, produced an interpreter and its own tests; Gary says the supplied example program ran.His hardest demonstration worked, but the model-authored tests were not an independent acceptance suite.

The progression matters. At 02:18, Gary spots the protocol example error. At 03:24, he moves to the interpreter task, which had caused older local models to loop, crash, or leak memory in his previous videos. This time he says the project finished in a few hours with planning and back-and-forth direction. The result is encouraging precisely because it was not a magical one-shot run.

Is It the New King of Local Models?

It may be an excellent option for a developer with suitable hardware, but this video alone cannot rank the whole field. We do not get a matched comparison with other current models on the same prompts, quantization, harness, machine, time budget, and independent test set. The author's source files and complete run logs are not linked in the video description, so readers cannot audit every generated artifact from the episode.

Qwen publishes its own benchmark table, including coding and agentic evaluations. Those are useful context, but they are vendor-reported measurements under specified harnesses. A more reliable purchasing decision comes from your own repository and acceptance criteria: does the model finish the job, how much review does it need, and how often must a larger model rescue it?

Local execution can keep prompts and code on a machine you control, but only if the model server, coding harness, tools, logs, and updates are configured accordingly. OpenCode's provider documentation describes connecting to local OpenAI-compatible endpoints. Check the selected provider and any enabled network tools before sending proprietary work.

A Fair Test on Your Own PC

  1. Record the exact setup. Model file and quantization, runtime version, GPU and VRAM, context length, thinking settings, and OpenCode provider all affect the result.
  2. Start with a task you can verify. Use one small project and write independent acceptance tests before asking the model to implement it.
  3. Keep a correction log. Count prompts, manual edits, failed runs, tool errors, and rescues by another model. A fast token stream is not the same as fast delivery.
  4. Repeat and compare. Run the same task more than once, then compare finished quality, wall time, and review effort with the model you already use.

For a first download, start with the official weights and model card, then choose a clearly labeled quantization from a source you trust. The Qwen repository documents serving paths; the OpenCode provider guide explains how a local server can be attached to a coding agent. Check actual memory use before increasing context or running long jobs.

Turn This Into a Business Idea

A local coding model is not a business on its own. The opportunity is a repeated customer job where privacy, turnaround, or customization really matters. This prompt uses Gary's tests to help find that job, then asks for a small paid experiment instead of a hardware purchase.

Business idea prompt

Find a local-first service worth testing

Match your skills to a buyer's repeated work and validate the offer before building infrastructure.

Ready to copy

Video Chapters

TimeWhat to watchTimeWhat to watch
00:00Why local Qwen matters00:32Model size and hardware
01:12Reported RTX 5090 speed01:54Single-file game
02:18Protocol spec and example error02:56Implementation and tests
03:24Interpreter challenge04:22Working result and caveats

Sources and Useful Links

Common questions

Does Qwen3.8-27B require 32GB of VRAM?
No universal 32GB minimum is established by this video. Gary used a 32GB RTX 5090 and a 4-bit quantized model. Some smaller quantizations or partial CPU offload may run with less VRAM, but speed, usable context, and quality will change.
Did Gary run the full-precision model on an RTX 5090?
No. He describes the full-precision weights as roughly 55GB and reports running a 4-bit quantized version on the 32GB RTX 5090. His reported speed of about 96 tokens per second applies to that setup.
Did Qwen3.8-27B complete the coding work without human input?
No. Gary describes giving the model a plan and follow-up direction in OpenCode. The protocol specification also contained small errors in worked examples before he proceeded to implementation and testing.
Is Qwen3.8-27B the best local coding model?
The video is a promising creator test, not a controlled ranking. Repeated tasks, independent tests, matched hardware, model files, context settings, and total completion time are needed to make that claim.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call