Local AI

M5 Ultra vs M3 Ultra for Local AI: What Got Faster?

Apple says the new Mac Studio can process LLM prompts up to four times faster than an M3 Ultra. NetworkChuck had an M3 Ultra on his desk and a loaned M5 Ultra to test that claim. His result is more useful than a single speed multiplier: the new machine reads a long prompt nearly four times faster, but writes the answer about one and a half times faster. He also tries the work people might actually do on a local AI computer: transcription, footage analysis, image and video generation, decision routing, agents, and video editing.

The short answer

In the creator's published five-run medians, Qwen3.8 27B at a 32,000-token prompt took 83.6 seconds to first token on M3 Ultra and 21.3 seconds on M5 Ultra. Writing then ran at 27.8 and 42.6 tokens/second, respectively. That is a big upgrade for document-heavy work, not a blanket 4x claim for every task.

Watch the M5 Ultra Local AI Test

Credit and disclosure: NetworkChuck's video was published 23 September 2026. Apple loaned him the M5 Ultra and he owns the M3 Ultra; he says Apple did not control his results. Micro Center sponsored the video, and its description contains affiliate links. This article is not sponsored. Measurements below are the creator's tests, not independent lab results.

What Was Actually Compared?

Both machines have 80 GPU cores, but this was not a matched-memory comparison: the M3 Ultra has 512GB unified memory and the M5 Ultra loaner 256GB. The creator used the same model, prompt, inference settings, and MLX versions for the chat tests. His method notes and raw data disclose different macOS versions and a stuck background process holding one M3 CPU core during most tests. The chat results use a warm-up and five measured runs; many other demonstrations are single runs. These details matter when you apply the numbers to your own machine.

Apple's launch announcement says M5 Ultra has Neural Accelerators in each GPU core and 1.2TB/s of memory bandwidth; M3 Ultra's published bandwidth is 819GB/s. Faster matrix computation plausibly explains the prompt-reading gain, while bandwidth plausibly explains the smaller writing gain. NetworkChuck did not profile the chip's bottlenecks directly, so this is an interpretation consistent with the data, not a measured causal breakdown.

The Benchmarks Without the Blur

WorkloadM3 UltraM5 UltraWhat to take from it
Qwen3.8 27B, 32K prompt: time to first token83.6 s21.3 sAbout 3.9x shorter wait; five-run median.
Qwen3.8 27B: answer generation27.8 tok/s42.6 tok/sAbout 1.5x faster writing, not 4x.
FLUX.2 klein 4B, 1024px image7.00 s1.77 sAbout 4x faster in the published image test.
Whisper large-v3 turbo, 2h 5m audio57.8 s27.3 sAbout 2.1x faster; speed and transcription quality are separate questions.
Qwen3-VL 32B, sampled 60s video61.4 s23.5 sAbout 2.6x faster at analyzing sampled frames, not full-video understanding.

These are figures from the creator's results page. The on-camera race displays slightly different single-run numbers: 82.8 versus 21.7 seconds to first token, and 28.1 versus 43.3 tokens/second for writing. Keep a live run separate from the published medians. Dense models showed larger long-prompt reading gains than the mixture-of-experts models in this test; model architecture and context length change the answer.

Seven Local AI Workflows Beyond Chat

  1. Transcription: Whisper on MLX processes a two-hour shoot in under two minutes with large-v3 in the video. The turbo model is quicker, with a quality trade-off to check on your audio.
  2. Footage descriptions: A local vision model samples video frames and produces a scene description. Chuck currently pays a cloud video API for some of this work; the demo shows feasibility, not proven equivalent accuracy.
  3. Images: mflux runs FLUX.2 locally. His published image benchmark has the largest gain in the table, but prompts, resolution, and quality settings will change runtime.
  4. Video generation: The LTX-2.5 demonstration produces short clips locally. The published 2-second and 4-second tests show different speedups, so there is no single video-generation multiplier.
  5. An agent brain: Hermes Agent can call a local model. Chuck gets it running but does not present the local model as his preferred replacement for every cloud-agent task.
  6. Fast decisions: He gives the small Laya model 300 labelled spam and banking questions and reports all 300 correct. That narrow test is not evidence it can safely approve real financial or security actions. Use confidence thresholds and human review.
  7. Editing software: A local model uses LM Studio and DaVinci Resolve's MCP interface to make a video edit. The on-camera workflow takes around 13 minutes and needs review; a successful tool call is not the same as an approved final cut.

Once models and software are downloaded, much of the inference can run without internet. That does not make every integration, download, update, or agent action offline. Nor does local processing remove the need for permissions, backups, data handling, and output checks.

Who Should Consider the M5 Ultra?

It is most compelling for frequent long-context inference, local video or image pipelines, and teams that can keep a high-memory machine busy while retaining control of their data. The M3 Ultra is still useful, especially when its 512GB capacity matters more than latency. A discrete GPU or rented compute may be a better value for smaller models or occasional jobs. For a larger comparison, read our DGX Spark test breakdown and cloud-versus-home cost guide.

Chuck discusses roughly $14,000 for his particular 256GB-memory, 8TB-storage M5 Ultra loaner. That is not a base price or a current quote. Before buying, run your own model, context length, and concurrent workload; measure time and cost per accepted result, expected utilization, power, software maintenance, and how often your work truly requires local processing.

Micro Center sponsorship in the source video

NetworkChuck points viewers to Micro Center's local AI offer, its Apple shop, and an Austin store-opening promotion. These are the creator's sponsored links, not recommendations or offers verified by this article. Check availability, terms, and dates directly.

Turn This Into a Business Idea

The interesting opportunity is not necessarily selling computers. It may be a narrow, reviewed workflow for a buyer whose repetitive audio, video, or document work benefits from local processing. This prompt starts with the buyer and a manual test, before any hardware purchase.

Business idea prompt

Find the job before buying the Mac

Compare a manual service, your current hardware, cloud compute, and a dedicated local machine.

Ready to copy

Video Chapters

TimeWhat to watch
00:00M5 Ultra reveal, loan, and sponsorship disclosure
01:25Hardware and memory bandwidth
01:57Qwen3.8-27B long-prompt race
06:54Model sizes and dense versus MoE
08:38Whisper transcription
10:25Vision model and footage
11:57FLUX.2 image generation
13:03LTX-2.5 video generation
14:18Hermes on a local model
15:37Laya decision model
16:43DaVinci Resolve MCP edit
17:55Final verdict and cost discussion

Sources and Tools

Common questions

Is the M5 Ultra four times faster than the M3 Ultra at all local AI tasks?
No. NetworkChuck measured about 3.9x faster long-prompt reading and time to first token with Qwen3.8 27B, but about 1.5x faster answer generation. Other workloads varied.
Were the two Mac Studios identical apart from the chip?
No. The M3 Ultra had 512GB unified memory and the loaned M5 Ultra had 256GB. They also ran different macOS versions, and the creator noted a background CPU process on the M3 during most tests. The same model, prompt, settings, and MLX versions were used for the chat comparison.
Does running these models locally mean no internet is ever needed?
Inference can run offline once the model weights and software are installed. Downloads, remote services, integrations, updates, and some agent actions can still require a connection.
Can this test justify buying a $14,000 Mac Studio?
Not by itself. That approximate figure describes the specific 256GB/8TB loaner configuration discussed in the video. Buyers should benchmark their own models and compare utilization, accepted-output cost, privacy needs, and alternatives before purchasing.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call