Qwen3.8-27B is a 27-billion-parameter open-weight model that Gary Explains tested on his own PC. In a six-minute video, he moves past short reasoning questions and uses OpenCode to make a game, implement a binary protocol, and build an interpreter in C. The interesting result is not that a model wrote code. It is that he could keep the model on a local machine and still finish work that had stopped smaller local models in his earlier attempts.
Watch Gary's Qwen3.8-27B Test
Credit: Gary Explains' original video, published 18 August 2026. The observations below come from the supplied transcript and his demonstration. Model details are checked against the official Qwen model card as of 1 October 2026. Gary is also on X and GitHub.
What Actually Ran on His PC?
Qwen's release is a dense 27B model with text and vision support, a native 262,144-token context window, and an Apache 2.0 license. Those are model properties, not a promise that a consumer GPU can use the entire context at a useful speed. The official project repository documents local serving options and points to quantized variants.
Gary estimates that the 16-bit weights are roughly a 55GB download, which does not fit wholly in the RTX 5090's 32GB of VRAM. He instead loaded a 4-bit quantization and measured about 96 tokens per second. That figure is his one-machine observation, not an official Qwen speed rating or an end-to-end coding time.
32GB is his tested setup, not a universal minimum. For example, a later Unsloth GGUF catalog lists a 4-bit variant around 17.5GB. The model file still needs headroom for runtime state, the KV cache, context, and other processes. A smaller card may require a different quantization, shorter context, or CPU offload, with different speed and quality. Do not infer Gary's 96-token result from file size alone.
What the Creator Tested
| Task | Observed result | What it tells us |
|---|---|---|
| Short reasoning questions | Answered the Alice and two-egg-timer questions. | A warm-up, not a demanding coding evaluation. |
| Single-file Space Invaders game | Produced an index.html with inline JavaScript that Gary opened and played. | It could deliver a small runnable artifact without external libraries in this test. |
| Type-length-value protocol | Drafted a specification, but one or two worked examples had incorrect numbers. | The design looked useful, yet its examples still needed verification. |
| Protocol library in OpenCode | Implemented the spec with unit and end-to-end tests after Gary supplied a plan and follow-up guidance. | The model completed a multi-step coding job with an operator in the loop. |
| New Scrippy interpreter in C | Inferred behavior from example code, produced an interpreter and its own tests; Gary says the supplied example program ran. | His hardest demonstration worked, but the model-authored tests were not an independent acceptance suite. |
The progression matters. At 02:18, Gary spots the protocol example error. At 03:24, he moves to the interpreter task, which had caused older local models to loop, crash, or leak memory in his previous videos. This time he says the project finished in a few hours with planning and back-and-forth direction. The result is encouraging precisely because it was not a magical one-shot run.
Is It the New King of Local Models?
It may be an excellent option for a developer with suitable hardware, but this video alone cannot rank the whole field. We do not get a matched comparison with other current models on the same prompts, quantization, harness, machine, time budget, and independent test set. The author's source files and complete run logs are not linked in the video description, so readers cannot audit every generated artifact from the episode.
Qwen publishes its own benchmark table, including coding and agentic evaluations. Those are useful context, but they are vendor-reported measurements under specified harnesses. A more reliable purchasing decision comes from your own repository and acceptance criteria: does the model finish the job, how much review does it need, and how often must a larger model rescue it?
Local execution can keep prompts and code on a machine you control, but only if the model server, coding harness, tools, logs, and updates are configured accordingly. OpenCode's provider documentation describes connecting to local OpenAI-compatible endpoints. Check the selected provider and any enabled network tools before sending proprietary work.
A Fair Test on Your Own PC
- Record the exact setup. Model file and quantization, runtime version, GPU and VRAM, context length, thinking settings, and OpenCode provider all affect the result.
- Start with a task you can verify. Use one small project and write independent acceptance tests before asking the model to implement it.
- Keep a correction log. Count prompts, manual edits, failed runs, tool errors, and rescues by another model. A fast token stream is not the same as fast delivery.
- Repeat and compare. Run the same task more than once, then compare finished quality, wall time, and review effort with the model you already use.
For a first download, start with the official weights and model card, then choose a clearly labeled quantization from a source you trust. The Qwen repository documents serving paths; the OpenCode provider guide explains how a local server can be attached to a coding agent. Check actual memory use before increasing context or running long jobs.
Turn This Into a Business Idea
A local coding model is not a business on its own. The opportunity is a repeated customer job where privacy, turnaround, or customization really matters. This prompt uses Gary's tests to help find that job, then asks for a small paid experiment instead of a hardware purchase.
Find a local-first service worth testing
Match your skills to a buyer's repeated work and validate the offer before building infrastructure.
Video Chapters
| Time | What to watch | Time | What to watch |
|---|---|---|---|
| 00:00 | Why local Qwen matters | 00:32 | Model size and hardware |
| 01:12 | Reported RTX 5090 speed | 01:54 | Single-file game |
| 02:18 | Protocol spec and example error | 02:56 | Implementation and tests |
| 03:24 | Interpreter challenge | 04:22 | Working result and caveats |
Sources and Useful Links
- Creator: Gary Explains' video, X, and GitHub profile. The video description does not link the generated code from these tests.
- Official model: Qwen3.8-27B weights, license, and model card and Qwen3.8 project and local-use notes.
- Local setup: Unsloth's quantized GGUF options, OpenCode local-provider documentation, and RTX 5090 memory specification.