Google will tell you its seventh-generation Ironwood TPU delivers roughly 44% lower total
cost of ownership than a comparable GB200 configuration. What Google will not tell you, three months
after Ironwood’s general availability, is what it actually costs to rent one. That asymmetry — bold
efficiency claims paired with no public price — is the defining friction in the
custom AI silicon vs GPU rental decision every infrastructure buyer with a
meaningful inference bill is now forced to evaluate.
Custom silicon’s TCO advantage over merchant GPUs is real for a narrow, specific workload profile
— stable, high-volume, well-understood inference on models already ported to the vendor’s software
stack — and largely irrelevant or actively counterproductive everywhere else. The pricing opacity
surrounding TPU and Trainium is not incidental; it reflects a market where custom-silicon economics
only make sense at negotiated, hyperscaler-adjacent scale, while GPU rental’s public, competitive
pricing remains the rational default for the vast majority of buyers below that threshold.
This article compares Google TPU v7 Ironwood, AWS Trainium3, and Microsoft Maia 200 against
NVIDIA GPU rental on published economics, lock-in risk, and workload fit. For the GPU-side pricing
baseline, see our H200 vs B200 vs H100 cost-per-token analysis and current
GPU cloud pricing benchmarks.
Why every hyperscaler is building a hedge against NVIDIA pricing
Every major hyperscaler now runs a custom AI silicon program: Google’s TPU line reaches its
seventh generation (Ironwood) in 2026, AWS ships Trainium3 as its first 3nm accelerator, Microsoft’s
Maia 200 targets inference specifically, and Meta’s MTIA v2 continues ramping internally. None of
these programs are framed publicly as “replacing NVIDIA” — all four hyperscalers remain among
NVIDIA’s largest customers simultaneously — but the strategic logic is the same across all of them:
reduce dependence on a single supplier whose data-center gross margin exceeds 75%, and capture some
of that margin internally at hyperscaler-relevant volume.
The custom AI silicon vs GPU rental question therefore looks different depending
on which side of the hyperscaler boundary a buyer sits on. For Google, Amazon, Microsoft, and Meta
themselves, custom silicon is a capital-allocation decision made at a scale where even a 10–15%
efficiency gain saves billions annually. For an external enterprise buyer renting compute from one
of these providers, the question is narrower and more concrete: does renting TPU or Trainium
capacity instead of GPU capacity save money on your specific workload, after accounting for migration
cost and lock-in — and can you even get a straight answer on price to find out?
Custom ASIC-based AI server shipments are projected to reach 27.8% of the overall AI server market
in 2026, growing at roughly 44.6% annually — nearly triple the 16.1% growth rate projected for
merchant GPUs over the same period. That growth is concentrated almost entirely in hyperscaler-internal
inference workloads, not in external rental products competing directly against NVIDIA GPU cloud
offerings, which is the central asymmetry this article unpacks.
TPU v7 Ironwood, Trainium3, and Maia 200: what the spec sheets actually say
Comparing custom silicon specifications against NVIDIA GPUs is complicated by each vendor
optimizing for a different point in the training-versus-inference, memory-versus-compute design
space, and disclosing performance in different precision formats. Reading past the headline
efficiency claims to the underlying hardware clarifies which workloads each chip actually targets.
Google’s TPU v7 Ironwood is unambiguously an inference-first, memory-bandwidth-first design: 192GB
of HBM3e per chip at roughly 7.37 TB/s bandwidth, rated at 4,614 FP8 teraFLOPS, with pods scaling to
9,216 chips sharing 1.77 petabytes of memory over a three-dimensional torus interconnect. That memory
profile targets exactly the bottleneck large-model inference actually hits — streaming model weights
and a growing KV cache through the chip fast enough to keep generation latency low — more than it
targets raw FLOPS supremacy.
AWS Trainium3, the company’s first 3nm accelerator, trails Ironwood on paper — 2.52 PFLOPS FP8 per
chip against 144GB of HBM3e at roughly 4.9 TB/s — but targets a different workload profile: dense
transformers and mixture-of-experts training and inference, with an UltraServer configuration packing
up to 144 chips for 362 FP8 PFLOPs aggregate. Microsoft’s Maia 200 takes a third approach entirely,
running scale-up over standard Ethernet rather than a proprietary interconnect, with 2.8 TB/s of
bidirectional bandwidth per accelerator and clusters scaling to 6,144 accelerators — but it is
inference-only, making a direct train-and-serve comparison against Ironwood or Trainium3 incomplete
by design.
Reported by RCR Wireless (2026): Microsoft claims Maia 200 delivers 3x the FP4 performance of AWS’s Trainium3, FP8 throughput above Google’s seventh-generation TPU, and 30% better performance-per-dollar than Microsoft’s own existing fleet — vendor-reported figures without full published test configurations, deployed first near Des Moines, Iowa, with Phoenix, Arizona next.
| Chip | Memory | Bandwidth | Peak FP8 | Workload focus | Rentable externally? |
|---|---|---|---|---|---|
| Google TPU v7 Ironwood | 192GB HBM3e | ~7.37 TB/s | 4,614 TFLOPS | Inference at pod scale | Google Cloud only; no public rate |
| AWS Trainium3 | 144GB HBM3e | ~4.9 TB/s | 2.52 PFLOPS | Training + inference | AWS EC2 only; limited public benchmarks |
| Microsoft Maia 200 | Not fully disclosed | 2.8 TB/s (per accelerator) | Above TPU v7 (vendor claim) | Inference only | Primarily internal/Azure OpenAI Service |
| NVIDIA B200 (reference) | 192GB HBM3e | ~8 TB/s | ~2.25 PFLOPS dense FP8 | Training + inference, general purpose | Yes — multi-cloud, neocloud, and spot markets |
Source: RCR Wireless hyperscaler chip comparison; TECHi — Ironwood specifications.
Ironwood optimizes for memory-bound inference at pod scale; Trainium3 targets balanced train-and-serve workloads; Maia 200 is inference-only — match the chip’s design target to your workload before comparing FLOPS figures directly.
Custom AI silicon vs GPU rental: the pricing-transparency gap
The single most consequential difference between custom silicon and GPU rental is not a
performance number — it is that NVIDIA GPU pricing is public, competitive, and shoppable across
dozens of providers, while TPU and Trainium pricing at the rates that make the TCO claims credible
is negotiated and disclosed selectively, if at all. A buyer evaluating B200 capacity in July 2026
could pull live marketplace data: $9.36 per GPU-hour on-demand, roughly $5.34–5.37 on spot, working
out to approximately $2.60 per million tokens on-demand or $1.48 on spot for a representative
Llama-70B FP8 workload. No equivalent on-demand figure exists for TPU v7 Ironwood three months after
its general availability.
That asymmetry matters practically, not just as a transparency complaint. A mid-size team serving
500 million tokens a day can build a defensible cost model on the NVIDIA side — roughly $1,300/day
before optimization at the on-demand rate above — and pressure-test it against revenue immediately.
The same team evaluating Ironwood has no equivalent number to enter into that spreadsheet; only a
contract rate reported for a customer operating at a scale thousands of times larger. The decision,
absent a comparable price point, defaults to the platform that actually answered the pricing
question.
Independent analysis attempting to reconstruct Ironwood’s implied economics from Google’s own TCO
claims estimated a cost of roughly $0.18 per million tokens for Gemini inference, against
approximately $0.31 per million tokens on comparable B200 configurations — consistent with the
headline 44% TCO advantage Google has cited. That comparison, however, reflects Google’s own
internal, fully-optimized deployment running its own model on its own infrastructure — not a rate
any external buyer can currently obtain by signing a standard cloud contract.
Reported by Spheron Network (July 7, 2026): NVIDIA B200 SXM6 on-demand pricing stood at $9.36 per GPU-hour with spot rates around $5.34–5.37, pulled from live marketplace data — working out to roughly $2.60 per million tokens on-demand for a representative Llama-70B FP8 workload. Ironwood would need approximately 171 tokens/second per chip just to match that on-demand rate, with no published number confirming whether it clears that bar.
Reported by Data Storage (2026): Independent analysis citing SemiAnalysis benchmarks estimated Ironwood’s cost at roughly $0.18 per million tokens for Gemini inference versus approximately $0.31 per million tokens on comparable B200 configurations — a comparison reflecting Google’s internal, fully optimized deployment rather than a rate available to external cloud customers.
Software lock-in: the migration tax GPU buyers don’t pay
A workload built on CUDA can typically move between AWS, Azure, Google Cloud, CoreWeave, and a
long tail of neoclouds with moderate engineering friction, because CUDA is the common substrate
nearly every training and inference framework targets first. A workload rebuilt for Google’s JAX/XLA
stack to run on TPU, or ported to AWS’s Neuron SDK to run on Trainium, loses that portability
entirely — there is no multi-cloud rental market for either chip, because neither is sold or hosted
outside its parent hyperscaler’s own cloud.
This is the least visible cost in any custom AI silicon vs GPU rental comparison,
and the one vendors have the least incentive to highlight. Teams that port a production inference
stack to Neuron or JAX to capture a genuine cost advantage today are simultaneously forfeiting the
negotiating leverage that comes from being able to credibly threaten to move workloads to a
competing GPU cloud provider — the exact leverage our
enterprise GPU procurement playbook identifies as a core protection
against provider price increases and service degradation.
The migration cost is not merely theoretical engineering time — it includes revalidating model
accuracy after a framework port, rebuilding CI/CD and monitoring tooling around a new SDK, and
retraining engineering staff on a stack with a materially smaller community and third-party tooling
ecosystem than CUDA’s. For teams running standard, widely-supported model architectures on stable
frameworks like vLLM, that migration tax frequently exceeds the marginal per-token savings a custom
chip promises, particularly once the pricing opacity discussed above is factored into the risk of
the decision.
When the economics do work: the Midjourney case and its limits
The clearest publicly reported example of custom silicon economics paying off is Midjourney’s
reported migration from NVIDIA GPUs to Google TPUs, which reportedly cut the company’s monthly
compute costs from $2.1 million to $700,000 — a 65% reduction. That case study appears throughout
custom-silicon marketing for good reason: it is a genuine, large, verifiable-scale example of the
TCO advantage materializing in production, not merely in a vendor benchmark.
The case also illustrates exactly the conditions under which the advantage appears: Midjourney
runs a narrow, well-understood, extremely high-volume inference workload (image generation from a
relatively stable model family) — precisely the profile that plays to TPU’s memory-bandwidth-optimized
design and justifies the engineering investment in a JAX-based serving stack. A team running a
diverse portfolio of frequently-changing models, or a workload still in active architecture
experimentation, would not capture the same result, because the migration and re-optimization cost
scales with model diversity and change frequency, not just raw token volume.
Enterprise buyers evaluating a similar move should treat Midjourney’s result as an existence proof
that the economics can work at sufficient scale and workload stability — not as a representative
outcome to expect by default. The 65% figure reflects a specific, well-suited workload after full
optimization; unoptimized or poorly-matched migrations regularly underperform their GPU baseline
instead.
Editorial estimate — Illustrative breakeven: migrating a stable inference workload from GPU rental to TPU. Methodology: Illustrative scenario for a team spending $200,000/month on B200 on-demand inference for a single stable, high-volume model, assuming a mid-range 35% cost reduction if fully migrated to TPU (below Midjourney’s reported 65%, reflecting a less ideal workload match), against an estimated 8–12 weeks of engineering time at a fully loaded cost of $20,000/week to port the serving stack to JAX/XLA and revalidate model accuracy. This is NOT a guaranteed outcome — actual savings depend entirely on workload stability and how well it matches TPU’s memory-bandwidth-optimized design.
| Line item | Estimate |
|---|---|
| Current monthly GPU rental spend | $200,000 |
| Projected monthly savings if migrated (35%) | $70,000 |
| One-time migration engineering cost (10 weeks) | ~$200,000 |
| Breakeven period | ~3 months |
Source: Editorial estimate — methodology stated above; substitute your own spend, model
stability, and negotiated TPU/Trainium rate before making a migration decision.
Worked example: below roughly $50,000/month in stable inference spend on a single well-understood
model, the migration engineering cost typically exceeds a full year of projected savings — below that
threshold, custom silicon migration is rarely worth the lock-in risk regardless of the headline TCO
percentage a vendor cites.
Reading the market-share numbers without overreacting to them
Headlines projecting NVIDIA’s inference market share falling from over 90% today to 20–30% by
2028 circulate widely in coverage of custom silicon’s rise, and they deserve scrutiny rather than
uncritical repetition. Estimates of NVIDIA’s overall AI accelerator revenue share for 2026 cluster
in the 75–85% range depending on methodology — IDC puts it closer to 81%, while a narrower
merchant-only comparison against AMD and Intel alone puts it above 87% — with custom silicon and
AMD together accounting for most of the remainder. Any of these readings sits comfortably inside a
“high-70s to high-80s” band; NVIDIA’s absolute data-center revenue keeps growing sharply regardless
of which end of that band is closest to reality, because the total addressable market is expanding
faster than NVIDIA’s share is moving in either direction.
Custom silicon’s 10–15% revenue share (and higher unit-shipment share, given lower average selling
prices than NVIDIA GPUs) is concentrated overwhelmingly in hyperscaler-internal workloads — the four
companies building this silicon deploying it against their own inference traffic — rather than
displacing external GPU rental demand from third-party buyers. The distinction matters directly for
this article’s audience: the market-share shift underway is evidence that custom silicon works at
hyperscaler scale for hyperscaler-controlled workloads, not evidence that an external enterprise
buyer should expect the same economics renting TPU or Trainium capacity as a customer.
Reported by Axis Intelligence (2026): Custom ASIC-based AI server shipments are projected to reach 27.8% of the overall AI server market in 2026, growing 44.6% year-over-year against 16.1% projected growth for merchant GPUs; Broadcom, the primary ASIC co-design partner, controls roughly 70% of that market with a reported backlog exceeding $30 billion.
The counterargument: shouldn’t falling GPU margins force NVIDIA to compete on price?
A reasonable objection to this article’s framing is that custom silicon’s growth should pressure
NVIDIA into more competitive pricing for external GPU rental, narrowing the TCO gap this article
describes rather than leaving it stable indefinitely. There is some evidence for this: aggressive
capacity build-out across neoclouds has already compressed GPU rental margins at the provider level,
which is part of why B200 spot pricing sits well below on-demand rates in competitive markets.
The response is that NVIDIA’s pricing power is protected by a different mechanism than pure
supply-demand balance: CUDA’s software moat means most buyers face a switching cost to leave the
NVIDIA ecosystem entirely, which custom silicon’s own steep migration tax (documented above)
reinforces rather than undermines. NVIDIA does not need to match TPU’s per-token economics dollar for
dollar — it needs GPU rental’s total cost of adoption (hardware cost plus zero migration friction) to
remain competitive with custom silicon’s total cost of adoption (lower hardware cost plus substantial
migration friction), which is a much easier bar to clear and explains why NVIDIA’s revenue continues
growing despite share erosion.
What this comparison can’t settle
The most significant limitation in this analysis is that neither Google nor AWS has published
transparent, independently verified pricing for their latest custom silicon generation as of August
2026 — every cost figure favorable to TPU or Trainium in this article derives from vendor TCO claims,
third-party estimation, or a single widely-cited customer case study, not a database of comparable
transactions. Readers should weight NVIDIA’s publicly verifiable pricing more heavily than custom
silicon’s claimed advantages simply because one side of the comparison can actually be checked.
A secondary limitation is workload generalizability: the Midjourney case study, the most concrete
public evidence of custom-silicon savings, involves a single company running a specific, unusually
stable, high-volume image-generation workload. Nothing in the public record confirms similar results
generalize to text-generation LLM serving, multimodal workloads, or any workload with meaningfully
different traffic patterns — treat that case study as an existence proof, not a template.
Finally, this comparison cannot account for the roadmap risk on either side: NVIDIA’s Rubin
platform (see our NVIDIA Rubin vs AMD Helios coverage) and the next
generation of custom silicon (Trainium4, TPU 8i, Maia’s successor) are all targeting late 2026 to
2027 availability, and any of them could shift this comparison meaningfully before a migration
decision made today finishes paying back its engineering investment.
Recommendation by workload profile: when custom silicon actually wins
Rent GPUs (default) if: your workload portfolio is diverse, your models change
frequently, you value multi-cloud negotiating leverage, or your inference spend is below roughly
$50,000/month on any single stable model — the migration tax to custom silicon will not pay back on
a reasonable timeline.
Evaluate custom silicon if: you run a single, extremely stable, high-volume
model family already compatible with JAX or Neuron, your monthly spend on that specific workload
exceeds six figures, and you can negotiate a rate directly with the hyperscaler rather than relying
on public pricing that may not exist yet.
Do not migrate for training workloads unless you are a hyperscaler running
internal pretraining at a scale that justifies dedicated engineering teams for framework
optimization — CUDA’s tooling maturity for novel architectures remains the safer default through at
least 2027.
What each role should verify, roughly in order of consequence:
- CTO / infrastructure lead (highest impact): quantify the true migration cost
(engineering time, accuracy revalidation, tooling rebuild) before comparing headline TCO percentages
between custom silicon and GPU rental. - ML platform lead: confirm your specific model architecture and traffic pattern
match a stable, high-volume profile before assuming vendor TCO claims apply to your workload. - Procurement / finance: request a negotiated rate quote directly rather than
relying on published TCO percentages, since neither TPU nor Trainium publishes transparent on-demand
pricing as of August 2026. - DevOps / MLOps: preserve a GPU-based fallback path during any custom-silicon
pilot — the lock-in risk documented above makes an unhedged full migration the highest-risk option
on the table.
One number would change this entire analysis: the first independently verified, publicly available
on-demand price for TPU v7 Ironwood or Trainium3. Until it exists, every custom AI silicon
vs GPU rental comparison — including this one — remains partly an exercise in reading
vendor incentives rather than comparing verified market prices.
FAQ: edge cases for silicon and procurement buyers
Can a startup actually rent TPU or Trainium capacity, or only large enterprises?
Yes — both Google Cloud TPU and AWS Trainium are available as standard cloud rental products, not exclusively negotiated enterprise deals. The catch is pricing transparency: smaller buyers pay published on-demand or committed-use rates while the steep discounts referenced in vendor case studies (like Midjourney’s reported 65% cost reduction moving to TPUs) typically reflect negotiated contracts at a scale most startups have not reached.
How risky is vendor lock-in if I build on TPU or Trainium instead of GPUs?
Materially higher than choosing among GPU cloud providers. A GPU workload built on CUDA can usually move between AWS, Azure, GCP, CoreWeave, and neoclouds with modest friction. A workload rebuilt on JAX for TPU or the Neuron SDK for Trainium is tied to a single cloud provider’s roadmap — there is no multi-cloud rental market for either chip.
Does the TCO advantage of custom silicon apply to training as well as inference?
The published advantages concentrate almost entirely on inference, particularly stable, high-volume, well-understood workloads. Training — especially frontier-scale pretraining with evolving architectures — still runs predominantly on GPUs because CUDA’s tooling and debugging ecosystem remains more mature for novel architectures than JAX/TPU or Neuron/Trainium.
Why doesn’t Google publish a public per-token or per-hour price for TPU v7 Ironwood?
Google has not stated a reason publicly, but the practical effect — whether intentional or not — is that only buyers already in a sales conversation can build a cost comparison against NVIDIA’s public GPU rental rates. Smaller teams evaluating a TPU migration should treat this pricing opacity itself as a decision-relevant signal, not just an inconvenience.
Is Microsoft Maia available for rent outside of Microsoft’s own workloads?
As of August 2026, Maia 200 is deployed primarily for Microsoft’s internal and Azure OpenAI Service inference workloads rather than as a broadly rentable instance type comparable to GPU rentals. Treat Maia as a signal of Microsoft’s cost-hedging strategy against NVIDIA pricing rather than a procurement option available to external buyers today.
What is the single biggest mistake teams make when evaluating custom silicon?
Comparing a vendor’s best-case, heavily-optimized benchmark number for their custom chip against an unoptimized, list-price GPU baseline. Every fair comparison requires matching optimization effort on both sides — an unoptimized TPU deployment can just as easily lose to a well-tuned GPU serving stack as the reverse.
Sources & further reading
- TECHi — Google’s Ironwood TPU still has no public price (2026)
- DEV Community — AWS Trainium3 vs NVIDIA H200 and B200 in 2026
- Spheron Network — Google TPU v7 Ironwood vs NVIDIA B200 inference cost
- RCR Wireless — Custom AI chips: how every hyperscaler’s silicon compares
- Data Storage — Google Cloud AI chip workload economics
- Algorithmine — NVIDIA Blackwell vs custom ASICs (2026)
- Axis Intelligence — AI chip market share 2026
- Celadon Research — NVIDIA AI accelerator market position, Q1 2026
- Silicon Analysts — NVIDIA AI GPU market share 2026
Related reading
- H200 vs B200 vs H100 cost per token — the GPU-side baseline this article’s custom-silicon comparisons are measured against.
- NVIDIA Rubin vs AMD Helios — how the next merchant-silicon generation raises the bar custom ASICs must clear.
- The enterprise GPU procurement playbook 2026 — contract and lock-in considerations that apply equally to custom-silicon rental commitments.
- GPU cloud pricing comparison 2026 — current market rates for the GPU rental side of this build-vs-rent-vs-custom decision.
Can a startup actually rent TPU or Trainium capacity, or only large enterprises?
Yes — both Google Cloud TPU and AWS Trainium are available as standard cloud rental products, not exclusively negotiated enterprise deals. The catch is pricing transparency: smaller buyers pay published on-demand or committed-use rates while the steep discounts referenced in vendor case studies (like Midjourney’s reported 65% cost reduction moving to TPUs) typically reflect negotiated contracts at a scale most startups have not reached.
How risky is vendor lock-in if I build on TPU or Trainium instead of GPUs?
Materially higher than choosing among GPU cloud providers. A GPU workload built on CUDA can usually move between AWS, Azure, GCP, CoreWeave, and neoclouds with modest friction. A workload rebuilt on JAX for TPU or the Neuron SDK for Trainium is tied to a single cloud provider’s roadmap — there is no multi-cloud rental market for either chip.
Does the TCO advantage of custom silicon apply to training as well as inference?
The published advantages concentrate almost entirely on inference, particularly stable, high-volume, well-understood workloads. Training — especially frontier-scale pretraining with evolving architectures — still runs predominantly on GPUs because CUDA’s tooling and debugging ecosystem remains more mature for novel architectures than JAX/TPU or Neuron/Trainium.
Why doesn’t Google publish a public per-token or per-hour price for TPU v7 Ironwood?
Google has not stated a reason publicly, but the practical effect — whether intentional or not — is that only buyers already in a sales conversation can build a cost comparison against NVIDIA’s public GPU rental rates. Smaller teams evaluating a TPU migration should treat this pricing opacity itself as a decision-relevant signal, not just an inconvenience.
Is Microsoft Maia available for rent outside of Microsoft’s own workloads?
As of August 2026, Maia 200 is deployed primarily for Microsoft’s internal and Azure OpenAI Service inference workloads rather than as a broadly rentable instance type comparable to GPU rentals. Treat Maia as a signal of Microsoft’s cost-hedging strategy against NVIDIA pricing rather than a procurement option available to external buyers today.
What is the single biggest mistake teams make when evaluating custom silicon?
Comparing a vendor’s best-case, heavily-optimized benchmark number for their custom chip against an unoptimized, list-price GPU baseline. Every fair comparison requires matching optimization effort on both sides — an unoptimized TPU deployment can just as easily lose to a well-tuned GPU serving stack as the reverse.