← Back to blog

The Local AI Stack on Apple Silicon: What Fits in 16, 32, 64, and 128GB

2026-08-25

By Vadym · Generated with Boba, curated by me


Listen to this article
TL;DR

On Apple Silicon in August 2026, your RAM tier decides your local stack more than your model choices do — 64GB holds a 27B LLM plus vision, speech-to-text, and TTS concurrently, and the work that still belongs in the cloud (multi-file coding, polished TTS, finished video) belongs there for quality reasons, not cost.

A 64GB Mac holds a 27B language model, a vision model, speech-to-text, and text-to-speech in memory at the same time — about 27GB total — with room left over to generate images. That is the actual story of local AI in August 2026: the binding constraint stopped being model quality and became RAM budgeting. Video generation is the one workload that still forces everything else out of memory. And the list of jobs that genuinely belong in the cloud has gotten short and specific.


The LLM Layer

The headline this week is Qwen3.8-27B. Alibaba released it while this post was being researched, and it supersedes Qwen3.6-27B as the best open-weight language model in the 27–30B class. Artificial Analysis's Intelligence Index puts it at 52, up from 38 for Qwen3.6-27B, and SWE-bench Pro climbs from 50.2% to 61.7%. Same ~18GB Q4 footprint. Same ollama pull qwen3.8:27b-mlx on Apple Silicon. Apache 2.0, no field-of-use restrictions. If you were running Qwen3.6-27B, update.

Muse Glimmer 30B (covered in the August 14 post) remains the secondary pick for narrow agentic and tool-use loops where its speculative-decoding speed matters more than factual reliability. Its 82% hallucination rate on knowledge tasks is the hard constraint — it is a model you route verified-data tasks to, not one you trust with open-ended factual questions. Qwen3.8-27B hallucinates at roughly 49%: about 2x Sonnet's rate, well below Glimmer's.

The coding gap has narrowed, but be careful about how you measure it. Qwen3.8-27B scores 61.7% on SWE-bench Pro; Sonnet 4.6 runs ~79% on SWE-bench Verified. Those are different benchmarks with different task sets, and Pro is the harder one, so subtracting them gives a number that means nothing. What you can say: the gap is smaller than it was in early 2026, and for high-volume, well-scoped tasks — structured extraction, routine tool calls, triage, monitoring — local is good enough. For multi-file reasoning, debugging complex systems, or anything where a wrong tool argument goes unnoticed downstream, frontier still earns its per-token cost.

On 16GB machines the 27B class does not fit. You are looking at 7B models — Qwen3-7B, Mistral 7B, Gemma3-4B — with meaningfully worse tool-use reliability (~52% on ReAct loops, against ~70% for the 27B+ class). Worth running for simple tasks, not for autonomous agent loops.

Image — Generation and Vision

Generation: FLUX.1-schnell via mflux is still the answer. Apache 2.0, commercial use included, no Llama-style field-of-use asterisk. Sub-15 seconds per image at 1024px on an M3 Max, 6–12GB depending on resolution and step count. The break-even against Midjourney or cloud APIs lands around 50+ images/week, but privacy is the sharper argument regardless of volume: personal photos, unreleased product designs, and client-confidential prompts should not go through a third-party API when a local option exists.

FLUX.1-dev and FLUX.1-Kontext (image editing and composition) offer higher fidelity but carry non-commercial licenses and take 4–6x longer. Worth knowing they exist; not the default.

Vision and OCR: Qwen3-VL leads OCRBench-V2 and is the most actively maintained VL family for MLX via mlx-vlm. The 4B variant runs in ~4GB, which makes it cheap enough to keep resident as a persistent "read this screenshot" tool inside an agent loop. The 32B variant at Q4 is ~18GB — the heavy-OCR option for dense tables and complex layouts, but on a 32GB machine it competes directly with your LLM for memory rather than sitting alongside it.

Audio — Transcription, Voice, and Music

Transcription: Whisper large-v3-turbo via mlx_whisper. At ~2GB and effectively free per audio hour, against $0.15–0.90/hr for Deepgram or AssemblyAI, this is the clearest local win in the stack on economics alone. Cloud STT only makes sense for enterprise diarization requirements or volume low enough that setup time dominates.

TTS: Qwen3-TTS via mlx-audio, ~3GB. Speaker similarity of 0.789 beats ElevenLabs v2's 0.646 on that metric, with cloning from a 3-second reference clip. Where ElevenLabs still wins: English prosody and emotional subtlety by ear, and latency — vendor-published figures put ElevenLabs Flash at 75ms first-chunk and Cartesia Sonic at 40ms, against ~120ms for Qwen3-TTS on M3-class hardware once conversion overhead is included. For interactive or polished output, cloud. For bulk, private, or always-on voice work, local.

Music and SFX: MusicGen via musicgen-mlx runs faster than real time on an M4 Max and is the only local music generation option with real MLX support. It is a prototype and background-audio tool. The gap to Suno or Udio is large for anything that needs to sound finished.

Video — Two Very Different Stories

Understanding: This is where local is genuinely strong. Qwen2.5-VL via mlx-vlm gives you native frame-rate sampling, quantized 7B and 32B checkpoints, and up to 1-hour video comprehension per Alibaba's documentation. The practical workflow is ~1fps frame sampling, which covers most agent-side use cases: reviewing screen recordings, processing meeting clips, reading dashboards. InternVideo2 and VideoLLaMA2 are effectively CUDA-only and do not belong in an Apple Silicon stack.

Generation: This changed this month. MiniMax H3 — Apache 2.0, open weights on Hugging Face — landed on August 3, and by August 10 Salvatore Sanfilippo (antirez, Redis creator) had shipped h3.c, a native Metal implementation for Apple Silicon. Reported numbers on an M5 Max: 512×512 clips in ~3.5 seconds at the fast 4-step preset, ~16.7 seconds at quality. The constraint is memory. The checkpoint is 33GB and peak RAM during generation hits ~40GB, which makes 64GB the floor rather than the comfortable tier — on a 64GB machine you are running H3 near the ceiling and everything else has to be unloaded first. 128GB runs it alongside the rest of the stack.

At 32GB the verdict is unchanged: local video generation is not worth the setup. Use Runway, Veo, or Kling. At 64GB+, MiniMax H3 is worth having for prototyping and iteration. Publication-quality output still goes to cloud.

What Your Mac Can Actually Run
Modality 16GB Mac 32GB Mac 64GB Mac 128GB Mac
LLM 7B models (Qwen3-7B, Mistral 7B) Qwen3.8-27B Q4 (~18GB) — primary pick Qwen3.8-27B + a second model resident 70B-class open weights at Q4 (~40GB), or Qwen3.8-27B with full headroom
Image gen FLUX.1-schnell (6–12GB) — alone, not with an LLM FLUX.1-schnell alongside the 27B LLM FLUX.1-dev / Kontext (slower, non-commercial) Any FLUX variant alongside the full stack
Vision/OCR Qwen3-VL-4B (~4GB) Qwen3-VL-4B alongside the LLM, or 32B Q4 (~18GB) alone Qwen3-VL-32B alongside the LLM 72B-class VL at Q4 (~40GB) for peak OCR
STT Whisper turbo (~2GB — fits any Mac) same same same
TTS Qwen3-TTS (~3GB — fits any Mac) same same same
Music/SFX MusicGen small MusicGen large MusicGen large, concurrent with the LLM Concurrent with everything
Video gen Skip — use cloud Skip — use cloud MiniMax H3 (~40GB peak, everything else unloaded) MiniMax H3 alongside the rest
Video understanding Qwen2.5-VL-7B Qwen2.5-VL-7B Qwen2.5-VL-32B Qwen2.5-VL-72B

A 64GB stack running concurrently: Qwen3.8-27B (~18GB) + Qwen3-VL-4B (~4GB) + Whisper turbo (~2GB) + Qwen3-TTS (~3GB) = ~27GB. That leaves ~37GB for FLUX.1-schnell or MusicGen to run without unloading the LLM. MiniMax H3 needs everything else to step aside.

Where Local Still Loses

Complex agentic coding. Sonnet and Opus retain a real edge on multi-file reasoning — holding a full codebase in context and generating correct patches. Local hallucination rates run 2–4x higher than frontier, which matters more on coding tasks, where a wrong function name or incorrect import goes straight into your repo.

TTS naturalness. By ear, ElevenLabs is better on English prosody and emotional delivery. The speaker-similarity metric favors Qwen3-TTS, but similarity to a reference clip and naturalness of the output are different things. For anything an audience hears, ElevenLabs Flash or Cartesia.

Video generation quality. MiniMax H3 producing 512×512 clips in 3.5 seconds is a real milestone. Runway and Veo producing 1080p+ clips with coherent motion and lighting is a different category.

Top-tier OCR at scale. Qwen2.5-VL-72B or InternVL3-78B scores need ~40GB for the VL model alone, which means a 64GB machine can run one — but only by itself.

Diarization and enterprise SLA. Deepgram and AssemblyAI have more mature speaker identification and production-grade reliability guarantees than any local pipeline today.


The split that falls out of this: local for high-volume, narrowly scoped, privacy-sensitive workloads where a downstream check catches errors. Cloud for multi-file coding, TTS an audience will hear, publication-grade video, and anything where wrong output goes undetected. Notice that none of the remaining cloud cases are about cost — they are all about quality or reliability. That is the shift 2026 actually delivered.


Sources

All benchmark figures come from the sources above. No sponsored content. I have not personally tested every model listed — inference hardware in use is an M3 Max 64GB.