AI DevelopmentCost Playbook11 min readPublished August 25, 2026

512GB unified memory · 1.2TB/s · the largest open-weight model still does not fit

M5 Ultra’s 512GB: Can a Desktop Hold a Frontier Model?

Apple announced the M5 Ultra Mac Studio today: up to 512GB of unified memory at 1.2TB/s, enough to physically hold most of the open-weight models checked here. Most — not all. The largest open-weight model in existence still does not fit, the 512GB tier has no published price, and bandwidth math still favors the datacenter. Here is the honest version of the capacity story.

DA
Digital Applied Team
Senior strategists · Published Aug 25, 2026
PublishedAug 25, 2026
Read time11 min
SourcesApple Newsroom · Unsloth · NVIDIA
M5 Ultra max memory
512GB
unified · 1.2TB/s
512GB tier: late Oct
Max orderable config
$18,299
256GB / 16TB SSD
512GB price unpublished
NVIDIA H200 bandwidth
4.8TB/s
NVIDIA preliminary spec
4x the Mac
Kimi K3 floor at ~1-bit
594GB
does not fit in 512GB

Apple’s M5 Ultra Mac Studio, announced August 25, 2026, offers up to 512GB of unified memory at 1.2TB/s — enough addressable memory to hold a frontier-class open-weight model for local inference on a desktop. That sentence has been the dream of the local-AI crowd for two years, and today it is roughly true. The interesting parts are the word “roughly,” the price Apple has not published, and the model that still does not fit.

The capacity threshold is real. DeepSeek-V4-Pro — 1.6 trillion total parameters — fits inside even the 256GB configuration you can order today. But Kimi K3, the single largest open-weight model in existence as of this writing, needs roughly 594GB even at 1-bit quantization. The headline claim fails on the actual largest model, and Apple's own flagship 512GB configuration is not orderable yet and carries no published price.

This post does one narrow job: it checks Apple's brand-new memory ceilings against the real disk footprints of four major open-weight models released before August 25, 2026, then runs the capex-vs-API break-even on published rates only. Capacity, bandwidth, and price — nothing else, because the rest of the local-AI buying decision is already covered elsewhere on this blog.

Key takeaways
  1. 01
    512GB at 1.2TB/s holds trillion-parameter-class weights.DeepSeek-V4-Pro (1.6T params, ~155GB quantized) and Kimi K2 (1T params, ~381GB) both fit inside the 512GB ceiling — trillion-parameter-class open weights, inside a desktop's memory.
  2. 02
    The largest open-weight model still does not fit.Kimi K3 (2.8T total parameters) needs roughly 594GB on disk even at its most aggressive ~1-bit quantization, per Unsloth's published tables. The capacity headline fails on the single biggest open model that exists.
  3. 03
    The 512GB flagship has no price.Apple says the 512GB configuration is coming in late October and has published no price for it. The maximum orderable configuration today is 256GB/16TB at $18,299; the M5 Ultra base is $5,499.
  4. 04
    Bandwidth cuts the other way.NVIDIA's H200 moves 4.8TB/s on NVIDIA's own preliminary specification — four times the M5 Ultra's 1.2TB/s. The Mac holds more memory than an H200's 141GB; datacenter silicon wins on bandwidth, which is what governs tokens per second at inference time.
  5. 05
    The break-even is ~13.9B tokens, with big exclusions.At DeepSeek-V4-Pro's published off-peak API rates, $18,299 of hardware equals roughly 13.9 billion blended tokens — before electricity, depreciation, and time-of-day API price swings, which the naive math ignores.

01The AnnouncementWhat Apple actually announced.

Apple’s newsroom release puts two chips in the new Mac Studio. The M5 Max carries an 18-core CPU, up to a 40-core GPU, and up to 128GB of unified memory at 614GB/s, from $2,499. The M5 Ultra is the headline: Apple’s first quad-die design, fusing two M5 Max processors over a next-generation UltraFusion interconnect into up to a 36-core CPU, up to an 80-core GPU with Neural Accelerators, a 32-core Neural Engine, and up to 512GB of unified memory at 1.2TB/s — which Apple describes as 50% more bandwidth than the previous generation. Pre-orders opened today in 30 countries; general availability begins September 22.

On AI, Apple’s silicon-focused companion release claims the M5 Ultra lets users “run huge LLMs with hundreds of billions of parameters entirely on device.” Treat that as what it is: vendor marketing copy. Apple attached no model name, no parameter count, and no tokens-per-second figure to the claim anywhere in either release. Whether the claim survives contact with real open-weight models is exactly what the rest of this post checks.

The same event refreshed the Mac mini around the 2-nanometer M6 from $899 — a capable desktop, not an inference machine, with memory bandwidth of 170GB/s on the 24GB and 32GB configurations and 153GB/s on the 16GB base. For how Apple's studio-class silicon stacked up against dedicated AI hardware before today, see our DGX Spark vs M5 Max vs RTX 6000 Blackwell comparison.

From $2,499
M5 Max
18-core CPU · up to 40-core GPU · up to 128GB @ up to 614GB/s

The volume configuration. 128GB of unified memory is enough for the efficient tier of open-weight models — DeepSeek-V4-Flash's recommended quant needs roughly 110GB all-in — but not for anything trillion-parameter class.

Ships Sep 22
From $5,499
M5 Ultra
up to 36-core CPU · up to 80-core GPU · up to 512GB @ 1.2TB/s

Apple's first quad-die chip — two M5 Max processors fused via next-gen UltraFusion. Base configuration is 96GB/1TB at $5,499; the 256GB/16TB maximum orderable build is $18,299. The 512GB tier arrives late October, price unpublished.

512GB: late October

One number that does not exist anywhere in this post: tokens per second. No independent throughput measurement of the M5 Max or M5 Ultra existed on August 25, because the hardware had not shipped. Several aggregator sites already circulate specific figures attributed vaguely to “prior-generation” hardware or unlabeled community runs, with no reproducible method disclosed. When real benchmarks land after September 22, they will settle the throughput question — until then, anyone quoting a tokens-per-second number for these chips is guessing.

02The Missing PriceThe flagship configuration has no price.

Here is the detail most launch-day coverage glossed over: the 512GB configuration — the one that makes the frontier-capacity argument possible at all — cannot be ordered today, and Apple has not said what it will cost. MacRumors confirmed that no price has been published for the tier. What you can order today tops out at 256GB of unified memory with 16TB of storage for $18,299.

"Mac Studio with 512GB of unified memory is coming in late October."— Apple Newsroom, August 25, 2026

That single unpriced line matters more than it looks. A capex-vs-API break-even is only as good as its capex number, and the exact configuration this launch’s headline argument depends on does not have one. Every piece of hardware math in this post therefore uses the two prices Apple has actually published — the $5,499 base and the $18,299 maximum orderable build — and treats the 512GB tier as what it is today: an announced capacity with an announced month and an unannounced cost. When Apple publishes the number in late October, the arithmetic below extends in five minutes. Until then, nobody can honestly tell you what frontier local inference at 512GB costs, because Apple has not said.

03The Capacity TestWhat actually fits in 128, 256, and 512GB.

The test is simple: take four major open-weight models released on or before August 25, 2026, look up the smallest practical quantization of each in Unsloth’s published quantization tables, and check them against Apple's three new memory ceilings. A model “fits” only if the weights load with enough headroom left for the KV cache, the OS, and a usable context window — for the quantization and KV-cache math behind these footprint numbers, see our standing reference.

Four major open-weight models released before August 25, 2026 — DeepSeek-V4-Flash, DeepSeek-V4-Pro, Kimi K2, and Kimi K3 — checked against Apple’s three new Mac Studio memory ceilings: M5 Max 128GB, M5 Ultra 256GB, and M5 Ultra 512GB. Columns: model, total and active parameters, smallest usable GGUF quantization, and whether the model fits at each memory tier. Quantization sizes are from Unsloth’s published model documentation.
ModelParams (total / active)Smallest usable GGUFM5 Max 128GB · from $2,499M5 Ultra 256GB · $18,299 max configM5 Ultra 512GB · unpriced, late Oct
DeepSeek-V4-Flash-0731unsloth.ai · DeepSeek-V4 docs284B / 13B103GB (UD-IQ3_XXS) — needs ~110GB RAM all-inYes — the only model checked here that fits the MaxYes — comfortableYes — comfortable
DeepSeek-V4-Pro-0813unsloth.ai · DeepSeek-V4 docs1.6T / 49B155.1GB near-lossless (UD-Q4_K_XL); 161.9GB lossless (UD-Q8_K_XL)NoYes — ~94GB of headroom even at the lossless quantYes — comfortable
Kimi K2 (2025)unsloth.ai · Kimi K2 run guide1T / 32B245GB floor (UD-TQ1_0, ~1.66-bit); 381GB recommended (UD-Q2_K_XL, 2.71-bit)NoNo — the 245GB floor leaves ~11GB, no room for KV cache and contextYes — the 381GB quant leaves ~130GB of headroom
Kimi K3 (weights: 2026-07-27)unsloth.ai · Kimi K3 docs2.8T / 104B594GB (UD-IQ1_S, ~1.66–1.92-bit dynamic) — needs ~610GB RAM+VRAM; full precision ~1.56TBNoNoNo — 82GB short even at the most aggressive quantization that exists

Read down the right-hand column and the announcement’s real shape appears. The 512GB tier holds three of the four models checked here, and that is why this machine matters. The 256GB configuration you can actually order today already holds DeepSeek-V4-Pro, a 1.6-trillion-parameter model, with roughly 94GB to spare. What no Apple memory tier holds is the last row.

04The Negative ResultThe largest open-weight model does not fit.

Kimi K3 is Moonshot AI’s 2.8-trillion-parameter sparse mixture-of-experts model — 896 experts, 16 active per token, 104 billion active parameters, native vision, a 1M-token context — and since its weights landed on July 27 it has been the single largest open-weight model in existence. Per Unsloth’s quantization tables, its smallest practical GGUF — a ~1.66-to-1.92-bit dynamic mix — is 594GB on disk and wants roughly 610GB of combined memory to run. Even shredded down to roughly one bit per weight, the model is 82GB too large for the biggest Mac Apple will sell this year.

This is where Apple’s “hundreds of billions of parameters” line deserves precision. Kimi K3’s active parameters — 104B — sit comfortably inside that phrase. But mixture-of-experts inference requires the full expert set resident in memory: the 2.8T-parameter total footprint is what has to fit, not the 104B slice that fires per token. By active-parameter count the vendor claim reads fine; by what actually determines whether a model loads, the largest open model overshoots the largest Mac by a comfortable margin. For a sharp independent read on what K3’s scale means for open weights generally, Interconnects’ analysis is the reference.

The honest scope
“The M5 Ultra can hold a frontier open-weight model” is true only for a subset of what frontier-class means in August 2026. DeepSeek-V4-Pro and Kimi K2: yes. Kimi K3, the largest open-weight model that exists: no — 594GB minimum against a 512GB ceiling. Open-weight scale grew past the biggest desktop in the same year the desktop caught up. Both sides of that race are still moving.

05The Trade-OffBandwidth cuts the other way.

Capacity decides whether a model runs at all; memory bandwidth largely decides how fast it generates once it does, because autoregressive decoding re-reads the active weights for every token. And on bandwidth, the datacenter comparison is not close. NVIDIA’s H200 moves 4.8TB/s through 141GB of HBM3e — four times the M5 Ultra’s 1.2TB/s, and nearly eight times the M5 Max’s 614GB/s. One caveat worth carrying: NVIDIA’s own product page still labels those H200 numbers “preliminary specifications.”

Memory bandwidth · Apple's new chips vs the datacenter anchor

Sources: Apple Newsroom; NVIDIA H200 product page, where the H200 figures are marked “preliminary specifications”
NVIDIA H200141GB HBM3e · datacenter, not purchasable at retail
4.8TB/s
M5 Ultraup to 512GB unified · from $5,499
1.2TB/s
M5 Maxup to 128GB unified · from $2,499
614GB/s

The two-way cut, stated plainly: the Mac wins on capacity — an H200 carries 141GB against the Mac’s 512GB ceiling — and datacenter silicon wins on bandwidth, which is what turns into tokens per second. That trade cannot be put on a per-dollar footing today: NVIDIA publishes no street price for the H200 (only OEM and cloud partners quote it, and it varies widely), and Apple has published no price for the 512GB tier. A model that merely fits is not the same as a model that runs at production speed.

How that trade-off resolves for the new chips is, for now, unknowable — the hardware ships September 22, and no independent measurement exists. What history suggests is direction, not magnitude: unified memory buys you models that datacenter cards need multi-GPU rigs to load, and costs you generation speed against those same rigs. Where the crossover sits for the M5 generation is precisely the benchmark to wait for.

06The EconomicsThe worked break-even, on published rates only.

Here is the arithmetic, built strictly from numbers that were published on August 25. The hardware side is the $18,299 maximum orderable Mac Studio — M5 Ultra, 256GB unified memory, 16TB SSD — because it is the priciest configuration with an actual price, and because it genuinely holds DeepSeek-V4-Pro, the largest open-weight model that both fits it and has a fully published metered rate. The API side is DeepSeek’s own price list: V4-Pro at $0.66 per million input tokens and $1.98 per million output tokens, off-peak, cache-miss.

Capex
Max orderable Mac Studio
$18,299

M5 Ultra, 36-core CPU, 80-core GPU, 256GB unified memory, 16TB SSD — the most expensive configuration Apple has actually priced. The 512GB tier has no published price and is excluded from every calculation here.

Apple / AppleInsider, Aug 25
Metered rate
Blended DeepSeek-V4-Pro rate
$1.32/M

Off-peak, cache-miss list price: $0.66/M input, $1.98/M output. At an illustrative 1:1 input-output mix that blends to $1.32 per million tokens. Peak-hour rates run double — the blend is the honest midpoint, not a promise.

api-docs.deepseek.com
Break-even
Hardware cost in token terms
~13.9B tokens

$18,299 ÷ $1.32 per million ≈ 13.9 billion blended tokens before the metered bill would have equaled the hardware spend. Illustrative arithmetic on published numbers — not a recommendation, and not a full TCO.

Naive break-even
What this math excludes
Deliberately absent from that 13.9-billion-token figure: electricity, the Mac’s depreciation and resale value, and time-of-day API price swings — DeepSeek’s published peak rates run twice its off-peak rates, so the metered side of the ledger moves with the clock. It also compares against the largest open-weight model that both fits the priced hardware and has a fully published metered rate: this is a DeepSeek-V4-Pro break-even, not a “frontier model” break-even, because the largest frontier open model does not fit any Mac you can price today. It is illustrative arithmetic, not a total cost of ownership.

Thirteen-point-nine billion tokens is a lot — and whether it is a lot for you is the entire question. A team running continuous multi-hour agent workloads can clear that volume in months; a team making occasional API calls never will. The deeper versions of this decision are already worked through in our local-vs-cloud subscription ROI analysis, the running-cost side in local workstation economics, and the fuller self-hosting picture in our frontier self-hosting TCO analysis — this post’s contribution is the new capacity threshold those analyses could not price, because it did not exist.

07The DecisionWho this machine is actually for.

Map the model you intend to run to the cheapest memory tier that holds it with headroom, and the buying decision mostly makes itself. Four honest lanes, from the quantization table above:

~103GB class
DeepSeek-V4-Flash and smaller

Fits the M5 Max's 128GB with runtime overhead accounted for. If your target model is Flash-class, the $2,499 Max is the value pick — the Ultra buys you capacity you will not use.

Pick M5 Max 128GB
~155–162GB class
DeepSeek-V4-Pro class

A 1.6T-parameter model, lossless-quantized, inside a machine you can order today with ~94GB of headroom. The strongest genuinely-available local frontier option this announcement creates.

Pick M5 Ultra 256GB
~350–380GB class
Kimi K2 and peers

Needs the 512GB tier at a workable 2.71-bit quant with ~130GB of headroom. That tier arrives late October at an unpublished price — if this is your lane, the honest move is to wait for the number before committing.

Wait for 512GB pricing
594GB+ class
Kimi K3

Does not fit any Apple memory tier at any quantization that exists. This lane stays on metered APIs or multi-GPU datacenter rigs regardless of what you think of the Mac Studio.

Stay on the API

Two forward-looking notes. First, the software side is not the bottleneck it was: Apple’s MLX stack is what this hardware is built to run, and our MLX framework guide covers the toolchain that turns unified memory into a working inference box. Second, expect the capacity race to keep moving under this machine: the largest open-weight total nearly tripled in roughly a year — Kimi K2 shipped at 1T parameters in 2025, K3 at 2.8T this July — and there is no reason to assume K3 stays the ceiling. A 512GB Mac that holds most of the models checked here may hold fewer of next year’s — which argues for buying against the model you will actually run, not against the frontier as a category. If you are weighing local inference against metered APIs for production workloads, our AI transformation engagements start with exactly this fit-and-economics evaluation on your own workload numbers.

08ConclusionA real threshold, honestly scoped.

The shape of local inference, August 2026

The desktop holds most of what we checked — and the frontier kept moving.

The M5 Ultra’s 512GB at 1.2TB/s is a real capacity threshold — trillion-parameter-class open weights, loading on a desktop. The 512GB tier itself arrives in late October — at a price Apple has not yet named. DeepSeek-V4-Pro fits the machine you can order today. Kimi K2 fits the machine you cannot order yet. Kimi K3 fits no Mac at all.

That last fact is the one to hold onto. The capacity headline is true for three of the four models checked here and false for the largest open-weight model that exists, the bandwidth comparison runs four-to-one against the Mac, and the break-even arithmetic — 13.9 billion blended tokens against the priced configuration — excludes enough real costs that it should start a decision process, not end one. Vendor capability claims with no model named and no throughput measured deserve exactly the scrutiny this post gave them.

The practical move is unchanged from every sound hardware decision: name the model you will actually run, price the memory tier that holds it with headroom, and run the token math on your own volumes. What this announcement changes is that for two of the four models checked here — DeepSeek-V4-Flash on the $2,499 Max and DeepSeek-V4-Pro on the $18,299 256GB unit — that calculation now has a priced desktop answer.

Get the local-vs-API decision right

Capacity math is checkable — run it before the invoice, not after.

Our team helps businesses evaluate local-vs-API inference economics on real workload numbers — model fit, hardware sizing, break-even math, and production rollout — delivered in days not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

Inference economics engagements

  • Model-to-hardware fit analysis on your target workloads
  • Capex-vs-metered break-even models with full exclusions stated
  • Open-weight deployment — quantization, MLX, serving stack
  • Multi-vendor API routing where local does not pencil
  • Cost & governance programs for hybrid local + API fleets
FAQ · M5 Ultra local inference

The questions we get every week.

Apple announced a new Mac Studio with two chip options: the M5 Max (18-core CPU, up to 40-core GPU, up to 128GB unified memory at 614GB/s, from $2,499) and the M5 Ultra (up to 36-core CPU, up to 80-core GPU with Neural Accelerators, up to 512GB unified memory at 1.2TB/s, from $5,499). The M5 Ultra is Apple's first quad-die chip, fusing two M5 Max processors via a next-generation UltraFusion interconnect. Pre-orders opened August 25 in 30 countries, with general availability from September 22. The 512GB memory configuration is the exception: Apple says it is coming in late October and has published no price for it.
Related dispatches

Continue exploring open-weight economics.