The most interesting part of Apple’s new M5 Ultra isn’t the usual “faster AI” claim. It’s having up to 512GB of unified memory with 1.2TB/s bandwidth in one desktop. The scale of what fits inside that memory is kind of absurd. A single Mac Studio can load models such as DeepSeek R1 671B, Kimi K2.6 1T and DeepSeek V4 Flash at usable quantizations. These are models we would normally associate with racks of GPUs, not a compact machine sitting under someone’s desk. This probably won’t change how the average person uses ChatGPT. But it could matter a lot for researchers, developers and companies that want powerful AI without uploading private data to someone else’s servers or paying for every token. The interesting shift is that self-hosting a massive model no longer automatically means building and maintaining a complicated multi-GPU system. Meanwhile, the M6 Mac mini tops out at 32GB, keeping it in the compact-model category despite its faster AI hardware. Local AI hardware seems to be splitting in two: affordable systems running increasingly capable small models, and high-memory workstations bringing previously server-class models into a single box. M5 Ultra model compatibility: https://canitrun.dev/gpus/m5-ultra/ Apple Silicon M1–M6 local LLM guide: https://canitrun.dev/guides/apple-silicon-llm-guide/ submitted by /u/MaySaki2
Originally posted by u/MaySaki2 on r/ArtificialInteligence
