Local AI hardware in the 128GB class costs more in 2026 because the memory is soldered and DRAM contracts jumped. This is a buy, wait, or rent decision for a machine you would use yourself. It is not a lab test, and it is not a spec teardown.
Bottom line – Memory is the product. NVIDIA raised the DGX Spark Founders Edition by 18%, from $3,999 to $4,699 (US, as of February 2026). A 128GB Strix Halo box sits about 75% above its pre-sale price (US, as of September 2026). – DRAM (dynamic random-access memory) relief is not expected soon. TrendForce’s July 2026 outlook, cited by DataHardware, says the supply gap widens through 2027. Waiting for a cheaper 128GB box is a weak bet. – If you need about 100GB or more today, for 100B-class mixture-of-experts (MoE) models, buy from shelf stock and choose by ecosystem. Use the Spark for CUDA (NVIDIA’s GPU software platform), Strix Halo for Linux and a lower price, or a Mac Studio for macOS and higher bandwidth. – If you run models of 32B parameters or less, you do not need 128GB. A 64GB box, or a graphics processing unit (GPU) with 24–32GB, costs far less. – If you use the machine only a few hours a week, renting a cloud GPU is cheaper. Use the break-even table below.
Last updated: October 3, 2026.
Why local AI hardware got expensive
You are paying for memory that is soldered to the board. On a Strix Halo mini PC, that DRAM is LPDDR5X, and you cannot add more later. DataHardware and OpenClawDC report that this memory is the largest cost of the finished box. When memory contracts rise, the shelf price rises with them.
TrendForce, cited by DataHardware, says conventional DRAM contract prices rose about 93–98% in the first quarter of 2026 versus the fourth quarter of 2025. The same report projected a further rise of about 58–63% in the second quarter of 2026.
NVIDIA says it raised the DGX Spark Founders Edition from $3,999 to $4,699 (US, as of February 2026) because of worldwide memory supply constraints. NVIDIA says the hardware did not change, and that the new manufacturer’s suggested retail price (MSRP) applies in every region. Orders already placed kept the old price. The change went live in the week of 23–27 February 2026.
Apple did the same on the previous Mac Studio. In March 2026, Apple removed the 512GB option on the M3 Ultra. Tom’s Hardware reports that Apple also raised the 256GB upgrade from $1,600 to $2,000 (US, as of March 2026).
If you are still choosing a full desktop, start with hardware for running powerful AI models locally. This is why waiting for the next model has not lowered prices in 2026.
| Product | Earlier price | Price in Sept-Oct 2026 | Change | Note |
|---|---|---|---|---|
| DGX Spark Founders Edition | $3,999 | $4,699 | +$700 (+18%) | NVIDIA MSRP change, Feb 2026 |
| GMKtec EVO-X2 128GB | about $1,999 (pre-sale) | $3,499.99 (1TB) | about +75% | street price, Sept 2026 |
| Minisforum MS-S1 MAX 128GB/2TB | $2,299 (launch) | $3,799 | +$1,500 (+65%) | Sept 2026 |
| Framework Desktop 128GB (DIY, no storage) | $1,999 (launch) | $3,449 | +$1,450 (+73%) | Sept 2026; out of stock/pre-order on 16 Sept |
| Mac Studio M3 Ultra 256GB upgrade | $1,600 upgrade | $2,000 upgrade | +$400 | March 2026, 512GB option removed |
Source: NVIDIA Developer Forums, Tom’s Hardware, OpenClawDC, Compute Market, Apple, as of October 3, 2026.
Price check: what each option costs in October 2026
Street prices here are US retail from September 2026, except the Ryzen AI Halo row, a July 2026 Micro Center listing. Bandwidth is in gigabytes per second (GB/s). Sources disagree on Framework prices between launch and September, so those in-between figures are left out. Prices and stock change weekly; check the seller before buying.
A search for dgx spark vs amd ai max 395 is a choice between NVIDIA’s CUDA box and a Strix Halo box built on the Ryzen AI Max+ 395. A gmktec evo-x2 vs dgx spark check is the same split at two street prices. An amd ryzen ai halo vs mac studio check compares a $3,999 developer kit (US, as of July 2026) with Apple’s M5 lineup.
OpenClawDC recorded several of these stickers on 16 September 2026. Compute Market’s list is dated 22 September 2026. Where those two pages disagree about earlier months, this article keeps only the September figures.
| Option | Memory | Memory bandwidth | Price (US, Sept-Oct 2026) | Availability note |
|---|---|---|---|---|
| NVIDIA DGX Spark (Founders Edition) | 128GB LPDDR5X | 273 GB/s | $4,699 | NVIDIA MSRP |
| GMKtec EVO-X2 (Strix Halo) | 128GB / 64GB | about 256 GB/s | $3,499.99 (128GB/1TB); $2,199.99 (64GB/1TB) | in stock 16 Sept; store notice says another increase is coming |
| Framework Desktop DIY (Strix Halo) | 128GB / 64GB | about 256 GB/s | $3,449 / $1,959 (without storage or OS) | out of stock / refundable pre-order on 16 Sept |
| Minisforum MS-S1 MAX (Strix Halo) | 128GB | about 256 GB/s | $3,799 (128GB/2TB) | in stock 16 Sept |
| AMD Ryzen AI Halo developer platform | 128GB | about 256 GB/s | $3,999 | Micro Center listing, July 2026 |
| Mac Studio, M5 Max | 36-128GB | up to 614 GB/s | from $2,499 (36GB base) | shipping since 22 Sept; 128GB configuration price was not in these sources |
| Mac Studio, M5 Ultra | 96GB / 256GB / 512GB | 1.2 TB/s | from $5,499 (96GB base) | 512GB ships “late October” per Apple; some reports say it may slip to January |
Source: NVIDIA, OpenClawDC, Compute Market, SpecPicks, Apple, Macworld, as of October 3, 2026.
All Strix Halo boxes use the same chip and memory, so inference speed is about the same across brands, and you pay for ports, cooling, support, and availability.
What performance do you actually get?
For large MoE models, the Spark, a 128GB Strix Halo box, and a Mac Studio are all usable. Large dense models are comfortable only when bandwidth is high. Prefill, the pass that reads your prompt, favors the DGX Spark.
An MoE model reads only the active parameters for each new token. It can stay quick near 270 GB/s: the Spark is rated at 273 GB/s, and Strix Halo at about 256 GB/s. A dense model reads every weight for every token. Bandwidth then caps how fast the reply arrives.
Prefill is mostly a compute job. Decode, the one-token-at-a-time reply, is a bandwidth job. Match both limits to the model you actually run before you pay for memory you will not use.
| Platform | Test | Decode (tokens/s) | Prefill (tokens/s) | Source and date |
|---|---|---|---|---|
| DGX Spark | gpt-oss-120b MXFP4, llama.cpp CUDA, short context | about 45 | about 2,050 | llama.cpp maintainer benchmark thread, March 2026 |
| DGX Spark | same, at 48k context depth | about 30 | about 1,225 | same thread |
| Strix Halo (128GB) | gpt-oss-120b, various backends/tools | about 31-57 (wide range) | about 340-720 | ServeTheHome, kyuz0 grid, AI Multiple; 2026 |
| Mac Studio M3 Ultra (256GB) | gpt-oss-120b | about 71 | not reported | AI Multiple |
| DGX Spark | Llama 3.1 70B dense, FP8, batch 1 | about 2.7 | not reported | LMSYS |
| Mac Studio M5 Max 128GB | Llama 3.3 70B Q4, MLX | about 15 | not reported | llmcheck.net (single source) |
Source: llama.cpp discussion, ServeTheHome, AI Multiple, kyuz0, LMSYS, llmcheck.net, as of October 3, 2026.
These figures come from different tools, quantizations, dates and settings. They are not directly comparable. Runtime choice alone has moved the same chip’s result by 2x or more. Treat them as ranges.
Decode reports on the 120B MoE model run from about 30 to about 71 tokens per second. Strix Halo prefill spans about 340 to about 720, so those write-ups do not agree. LMSYS reports about 2.7 tokens per second for dense Llama 3.1 70B in FP8 (8-bit floating point) on the Spark at batch 1. One site, llmcheck.net, reports about 15 tokens per second for Llama 3.3 70B at Q4 (4-bit) on a 128GB M5 Max using MLX (Apple’s machine-learning framework).
Apple says the M5 Ultra is up to 4x faster than the M3 Ultra at LLM (large language model) prompt processing in LM Studio. That is Apple’s claim. Independent M5 Ultra figures were not available in our sources as of October 3, 2026.
AI Multiple reports 124 tokens per second on the 120B model for three RTX 3090 cards (72GB total). That rate is higher than the boxes above. It adds power draw, noise, and setup effort. No GPU street price is quoted here.
Buy now, wait, right-size, or rent: a decision guide
Choose this local AI hardware by workload, not by brand.
If you searched for the “best unified memory pc”, start from the job. There is no single winner here. The Spark is the CUDA and prefill buy.
Strix Halo is the lower 128GB street price. The Mac Studio is the bandwidth buy. A rented GPU sells hours, not a chassis.
| If you … | Do this | Why |
|---|---|---|
| Need 100GB+ of memory today for 100B-class MoE models or client data that cannot leave your site | Buy from shelf stock now | Forecasts do not show DRAM relief before late 2027 (F6); one store already warns of another increase (F4). |
| Run models up to 32B parameters (4-bit) | Right-size: 64GB box ($1,959-$2,199.99 for Strix Halo, F4) or a 24-32GB GPU | The 128GB step-up is $1,300 on the EVO-X2 and $1,490 on the Framework Desktop, calculated from September 2026 prices. |
| Use it a few hours per week | Rent a cloud GPU | See break-even table below. |
| Want the CUDA stack and fast prompt processing | DGX Spark, accepting the highest price per GB | CUDA ecosystem; prefill about 2,000 tokens/s on gpt-oss-120b (F8). |
| Want the most bandwidth per dollar for dense 70B-class models | Mac Studio M5 Max/Ultra, once you verify configuration price and independent benchmarks | 614 GB/s to 1.2 TB/s vs 256-273 GB/s on the other boxes (F7). |
| Serve several users or run production | Do not use a desktop box; use rented or dedicated GPUs | See GPU VPS providers for AI and RunPod, Vast.ai, and Lambda Labs. |
Source: TrendForce via DataHardware, OpenClawDC, NVIDIA, llama.cpp discussion, Apple, as of October 3, 2026.
Buy from stock if you need about 100GB this month. TrendForce’s 30 July 2026 outlook, cited by DataHardware, says the DRAM gap widens through 2027. New fab capacity is unlikely in volume before late 2027 or 2028. OpenClawDC recorded a GMKtec notice of another price increase on 16 September 2026.
Right-size the box when the model is 32B or smaller at 4-bit. Calculated from those September prices, the Framework step from $1,959 (64GB) to $3,449 (128GB) is $1,490. The EVO-X2 step from $2,199.99 to $3,499.99 is $1,300. A 24–32GB GPU is the other small path, and this article does not price those cards.
Rent local AI hardware time when the box would sit idle most of the week. The hours are in the next table. One desktop is also one queue. For several users, use a rented or dedicated GPU.
The Spark costs more per gigabyte than the EVO-X2. Calculated from the $4,699 MSRP (US, as of February 2026) and 128GB, that is about $36.7. Calculated from the EVO-X2 price of $3,499.99 (US, as of September 2026) and 128GB, that is about $27.3. You pay the gap for CUDA and for prefill near 2,050 tokens per second in the March 2026 llama.cpp run.
The Mac is the bandwidth buy. Apple rates the M5 Max at up to 614 GB/s and the M5 Ultra at 1.2 TB/s. The Spark is 273 GB/s. Strix Halo is about 256 GB/s.
The M5 Ultra base is $5,499 (US, as of September 2026). MacRumors says that base is $200 above the post-June M3 Ultra price and $1,500 above the M3 Ultra launch price. The M3 Ultra is rated at 819 GB/s. Apple’s 4x prompt claim does not replace a third-party test.
Rent or buy: break-even math
Break-even hours equal the purchase price divided by the hourly rental rate. The hourly rates below are illustrative assumptions, not quotes. Check current prices on the GPU cloud pricing comparison.
| Purchase price | If rental is $1.00/h | If rental is $1.50/h | If rental is $2.50/h |
|---|---|---|---|
| $2,499 | 2,499 h (625 days) | 1,666 h (416 days) | 1,000 h (250 days) |
| $3,499 | 3,499 h (875 days) | 2,333 h (583 days) | 1,400 h (350 days) |
| $4,699 | 4,699 h (1,175 days) | 3,133 h (783 days) | 1,880 h (470 days) |
Source: calculated from the purchase prices in this article and illustrative rental rates of $1.00, $1.50, and $2.50 per hour, as of October 3, 2026.
The day counts assume four hours of use per day. This math ignores electricity, resale value, and the chance that a rented GPU has more memory bandwidth than the desktop box. Heavy daily use pushes the answer toward buying.
Is waiting worth it?
Waiting for a cheaper copy of today’s 128GB spec is not what the forecasts support. TrendForce’s July 2026 outlook says the DRAM gap widens through 2027. New fab output is unlikely in volume before late 2027 or 2028. GMKtec has already warned of another hike.
Wait only if you need more than 128GB and a named product is close. Apple says 512GB M5 Ultra models ship in late October 2026. Macworld reports that some of those orders may slip to late January. A 256GB unified memory PC in this lineup is the M5 Ultra, and the date is not firm until it ships.
OpenClawDC reported a Ryzen AI Max+ PRO 495 with 192GB on 16 September 2026. No price was given. That report is unconfirmed. Wait for it only if 128GB is too small.
Who should not buy a 128GB unified-memory box
Skip this local AI hardware when the model, the speed target, or the user count does not need it.
- You only run 7B–14B models. Those weights fit in much less memory, so the 128GB step-up calculated above buys capacity you will not use.
- You want a fast chat on a dense 70B model and you are not buying a high-bandwidth Mac. LMSYS puts the Spark near 2.7 tokens per second on Llama 3.1 70B. Strix Halo’s bus, about 256 GB/s, sits near the Spark, not near a Mac Ultra.
- You plan to fine-tune or train in a serious way. The figures in this article are inference tests. A real training job wants a different GPU stack than a soldered mini PC.
- You need to serve several users. One desktop is a single queue. Use a rented or dedicated GPU host instead.
Checklist before you pay
- Confirm the exact memory and storage, including 64GB versus 128GB and 1TB versus 2TB.
- Confirm the memory is soldered. On Strix Halo you cannot add more later.
- Check whether storage and an operating system are included. Framework’s DIY price, $1,959 or $3,449 (US, as of September 2026), includes neither.
- Compare US, EU, and UK stickers. LLMRequirements reports that in mid-May 2026 Framework raised the EU 128GB price 12%, from EUR 3,029 to EUR 3,379. The UK price rose 11%, from GBP 2,699 to GBP 2,999. The US price did not change then.
- Read the return rules before you pay a deposit. Framework’s 16 September listing was a refundable pre-order.
- Check stock and the ship date on the day you pay. OpenClawDC showed the EVO-X2 and the MS-S1 MAX in stock, and the Framework board out of stock.
- Confirm your software stack: CUDA, ROCm (AMD’s GPU software stack), or MLX and Metal on a Mac.
- Look for a price-increase notice on the seller’s page. GMKtec had one on 16 September 2026.
Bottom line
Buy local AI hardware in the 128GB class only if a current project needs that memory now and a seller can ship that exact configuration from stock. If your models stay at 32B parameters or below, buy a 64GB box or a smaller GPU, and if the machine will sit idle most days, rent a cloud GPU instead of paying the full sticker up front. For detailed specs and benchmark methodology, see our DGX Spark vs Ryzen AI Halo vs Mac Studio comparison.
Frequently asked questions
Is the DGX Spark worth $4,699?
It is worth $4,699 (US, as of February 2026) only if you need CUDA and that fast prefill. Calculated from the MSRP, the Spark is about $36.7 per gigabyte, and calculated from $3,499.99, the EVO-X2 128GB model is about $27.3 per gigabyte. You pay more per gigabyte and get NVIDIA’s software stack. If you do not need CUDA, buy a Strix Halo box when it is in stock.
Will DRAM prices fall in 2027?
TrendForce’s outlook of 30 July 2026, cited by DataHardware, says the supply-demand gap widens through 2027. That outlook also says new fab capacity is unlikely to arrive in volume before late 2027 or 2028. One forecast is not a consensus, and forecasts change. Nothing here guarantees a 2027 drop or a further rise, so plan on today’s stickers if you need the machine now.
Is a Mac Studio better than a DGX Spark for local LLMs?
A Mac Studio is the higher-bandwidth machine, which helps dense models. Apple rates the M5 Ultra at 1.2 TB/s, against 273 GB/s on the Spark. The Spark is the CUDA machine, with strong prefill in the March 2026 llama.cpp run. Independent M5 Ultra tests were not in these sources as of 3 October 2026, so pick by the limit you actually hit.
How much memory do I need for a 70B model?
You need space for the 4-bit weights and space for the context cache. A 128GB box is more machine than a modest-context 70B model requires. This page does not state a gigabyte total for that sum, because the price sources did not include it. Do not buy 128GB only because a 70B model sounds large.
Should I rent GPUs instead of buying?
Rent if you will not reach the break-even hours in the table above. At an illustrative $1.50 per hour, a $3,499 purchase takes 2,333 hours, or 583 days at four hours a day. Those rates are not live quotes. If you will run the box for many hours every day, buying gets closer, and if the calendar is sparse, keep the cash and rent the hours.
Sources
- NVIDIA Developer Forums, 23 February 2026 price announcement
- Tom’s Hardware, DGX Spark price increase
- NVIDIA Developer Forums, 273 GB/s thread
- DataHardware, local AI mini PC prices
- OpenClawDC, which Strix Halo mini PC to buy
- Compute Market, Strix Halo prices
- SpecPicks, Ryzen AI Max+ 395
- LLMRequirements, Framework EU and UK price change
- Apple Newsroom, Mac Studio with M5 Max and M5 Ultra
- MacRumors, M5 Max vs M5 Ultra
- Macworld, 2026 Mac Studio timing
- Tom’s Hardware, M3 Ultra 512GB option removed
- llama.cpp discussion 16578
- LocalAIMaster, DGX Spark vs Strix Halo vs Mac Studio
- AI Multiple, DGX Spark alternatives
- Compute Market benchmarks
Updated: October 2026. Prices are US retail figures from the named sources, with the dates given in the text. This page did not test the machines. The rental rates in the break-even table are assumptions, not quotes. This is not a bid and not a forecast that DRAM prices will fall.