You keep the human work. We build the rest.
If you run a service business, that line is the whole product. The part most people skip is the box. A private agent that never leaves the shop still has to load an open-weight model into unified memory. The DRAM shortage made that box a luxury. This is the map of what actually fits, from files we opened on 25 August 2026. No invented speed tests. No invented dollars.
Custom AI systems start with the file, not the billboard
Apple launched M6 and M5 Ultra today. The $899 Mac mini is not a local-70B machine. M6 unified memory tops out at 32 GB. Llama 4 Scout at 2.71-bit is already 42.2 GB of weights. The “hundreds of billions on device” line sits on the Mac Studio with M5 Ultra, up to 512 GB, from $5,499, with 512 GB arriving late October. Apple printed the date. It did not print the dollars.
That is why the 70B chart was the wrong chart. An owner-operator does not buy a billboard. They buy a private agent that can hold the work they named.
[Image blocked: Open-weight models that fit local AI unified memory]
Mac Studio and the rest of the hardware ladder
Mac mini, M6 and M5 Pro. M6: 16 / 24 / 32 GB, up to 170 GB/s, $899 (education $799). M5 Pro: 24 GB base at $1,699, CTO 48 or 64 GB, 307 GB/s. Apple did not print 32 GB as an M5 Pro step. Pre-order 25 August. Ships 22 September. (Newsroom, specs, buy)
Mac Studio, M5 Max and M5 Ultra. M5 Max: 36 GB base at $2,499, CTO 48 / 64 / 128 GB, 460 GB/s at the base and 614 GB/s with the 40-core GPU. M5 Ultra: 96 GB base at $5,499, CTO 256 or 512 GB, 1.2 TB/s. 512 GB is a real spec and a late-October SKU. List price is blank. (Newsroom, specs, buy)
MacBook Pro is M5, not M6. 14-inch M5 is 16 / 24 / 32 GB at $1,999 / $2,199 / $2,399. M5 Pro 24 / 48 / 64. M5 Max 36 / 48 / 64 / 128. 96 GB on a MacBook Pro is not printed. (specs)
Framework Desktop. 32 / 64 / 128 GB LPDDR5x, soldered. The only printed dollar we opened is the default 32 GB mainboard at $969. Complete-system totals did not print. 192 GB is not a current buy step. (frame.work/desktop)
NVIDIA DGX Spark. One RAM step: 128 GB. Marketplace $4,699. NVIDIA staff, 25 February 2026: MSRP “adjusted from $3,999 to $4,699 due to memory supply constraints.” (marketplace, forum)
I sit on 64, 128, and 512. Inventory. The floor we use when we stand up on-prem AI. Not a bench.
DRAM shortage: the printed prices moved
Apple raised the Mac floor from 8 GB to 16 GB in late 2024. The M4 mini launched at $599. Then 2026 cut the high-RAM SKUs. Then the list reset. On 25 June Apple said it had “never seen a component price increase this much, this quickly.” M3 Ultra went $3,999 to $5,299 on that snapshot. The M6 mini is $899. Same 16–32 GB ceiling.
NVIDIA had already moved Spark $3,999 to $4,699 in February. Same reason, in writing.
[Image blocked: DRAM shortage: what a local AI box costs now]
TrendForce, 1 June 2026: conventional DRAM contracts up about 93–98% quarter-on-quarter in Q1. Industry DRAM revenue +81% to $97 billion. They then expected another 58–63% in Q2. Incremental supply goes to high-capacity RDIMMs for AI servers. (TrendForce 1 Jun)
HBM is taking the wafers. Top three suppliers: about 18% of DRAM wafers into HBM at end-2025, 22% at end-2026, 30% at end-2027. Bits are smaller still, 8 / 9 / 13. (TrendForce 2 Jun)
SK hynix CEO Kwak Noh-jung, 10 July 2026: “next year will be the worst year in the industry’s history from the supply perspective.” Demand above their capacity “even beyond 2030.” (Reuters)
The wafers go to the cloud. A service-business owner who wants a private agent gets what is left, at the printed price.
Open-weight models that fit, by unified memory
Fit rule: opened GGUF or MLX file size, plus Unsloth’s own total-memory floor when they print one. OS and context memory are extra. No speed numbers.
[Image blocked: On-prem AI: open-weight file size vs unified memory]
16 GB. Gemma 4 E4B Q4 at 4.98 GB. Gemma 4 12B Q4 at 7.12 GB. Short chat, light tools. Cannot: Qwen3.8-27B 4-bit. Unsloth names a 24 GB Mac for that file.
24 GB. Qwen3.8-27B 4-bit at 16.1–16.8 GB. Everyday local chat, vision, coding. Still not Scout. Still not a 70B box.
32 GB. 27B Q6 is the comfortable pick. 27B Q8 at 29 GB is tight. Still not Scout Q2 at 42.2 GB.
48 GB. Scout Q2 is a squeeze. Coder-Next 4-bit at 45–49 GB only if you accept Unsloth’s “>45 GB” as empty-box math.
64 GB. A real local coding agent (Coder-Next) plus the OS. Scout Q3 at 52.9 GB. Cannot: Flash. Flash is 82.5 GB even at 1-bit.
96 GB (Mac Studio Ultra base). First step that can attempt DeepSeek-V4-Flash, at 1-bit (82.5 GB). Flash Q2 at 96.8 GB leaves no OS.
128 GB. Flash Q2. Flash Q4 at 155 GB does not fit. This is Spark, Framework 128, and M5 Max 128.
256 GB (Mac Studio). Flash Q4 at 155 GB and Flash Q8 at 162 GB. Unsloth calls Q8 near-lossless versus Q4. Cannot: Qwen 2.4T.
512 GB (Mac Studio, late October; also the 512 GB M3 Ultra we already have). Smallest opened Qwen 2.4T GGUF is 397 GB. Unsloth wants at least 450 GB RAM. 2.4T IQ1_S is 508 GB — no OS left. V4-Pro Q2 at 574 GB does not fit.
The gates: ~5 GB, ~16 GB, ~45 GB, ~83 GB, ~155 GB, ~397 GB. Not a single 70B cliff.
On-prem AI versus the cloud agent
OpenAI, 27 February 2026: ChatGPT has more than 900 million weekly users and more than 50 million consumer subscribers. (OpenAI)
No opened study prices a 512 GB Mac Studio against that subscription. We will not invent a payback.
Most people will stay on a hosted agent. That is fine for work that can leave the building. An owner-operator who needs private agents — patient records, member files, the book that does not go to someone else’s GPU — needs a box that holds the open-weight model, on-prem, with the human checkpoints they already named.
Apple sells the Studio as on-device models “without counting tokens.” NVIDIA sells Spark the same way. Those are vendor sentences, not a cost study. The useful question is still: does the file fit.
What Utlyze actually builds
We do not sell you a Mac and walk away. We name the work that stays a person. We build the custom AI system that does the rest, on local AI that can hold the model.
Dental recovery. Membership retention. The outreach that has an approved review point. The map starts with five questions on utlyze.com. The hardware is the floor, not the product.
If the $899 mini is the box in the shop, you get Gemma / 27B-class agents. If the work needs Flash-class or a 2.4T attempt, you are on Mac Studio unified memory, and you are paying luxury-box money because DRAM went to HBM first.
Name the work. Book a call. We build the rest.