Meta's 'Iris' AI Chip Enters Production in September — Why a $125–145 Billion Spender Wants Off Nvidia

Meta's 'Iris' AI Chip Enters Production in September — Why a $125–145 Billion Spender Wants Off Nvidia
An internal Meta memo reviewed by Reuters says the company will start manufacturing its own AI chip — codenamed "Iris," part of the MTIA program — in September 2026, aiming to roughly double its computing capacity to 14 gigawatts by next year. One test chip reportedly passed validation in about six weeks with no major issues. It's a direct bid to lower a GPU bill buried inside a $125–145 billion capital-spending budget. This piece explains what Iris actually is, why Meta is building it now, and whether an in-house chip can realistically dent Nvidia's grip.

Meta is the latest hyperscaler to decide that renting compute from Nvidia forever is too expensive to accept. According to a memo reviewed by Reuters, its in-house Meta Training and Inference Accelerator (MTIA) program will move a chip codenamed Iris into production this September. The goal isn't to beat Nvidia at making the world's best general-purpose GPU — it's to build a chip tuned so tightly to Meta's own workloads that it's cheaper to run at Meta's scale. That distinction is the entire story, and it's why this move matters even though Meta will keep buying Nvidia silicon for years.

What Iris actually is

Iris isn't a science project. Here's what the reporting establishes.

Detail Specifics
Program Meta Training and Inference Accelerator (MTIA), in-house
Chip codename Iris
Production start September 2026
Compute target ~7 GW planned in 2026 → ~14 GW in 2027 (double)
Test-to-validation time ~6 weeks, no major issues reported
Design partner Broadcom
Manufacturing TSMC
Supporting suppliers Samsung (RAM), SanDisk (storage), Sumitomo Electric (fiber optics)
Primary jobs Training ranking/recommendation models; broad inference
Stated goal Lower GPU costs; reduce Nvidia dependence

Two numbers carry the weight here. The six-week validation signals the design is maturing fast — Meta isn't stuck in silicon purgatory the way first-time chip efforts often are. And doubling capacity to ~14 gigawatts is a data-center build-out on the scale of a national power grid, which is exactly why shaving even a slice of per-chip cost matters so much.

Concept diagram of Meta's Iris AI chip supply chain with Broadcom, TSMC, Samsung and SanDisk as partners

## Why now: the math behind "get off Nvidia"

Meta's 2026 capital expenditure is guided at a staggering $125–145 billion, and a large chunk of that historically flows to Nvidia (and AMD) GPUs. Meta has said it "still expects to spend plenty" on those — so Iris isn't a clean break. It's a hedge against a single supplier owning the most expensive line item in the company's budget.

The logic mirrors what other giants have already concluded. A custom chip only has to be good enough for one company's specific tasks — in Meta's case, ranking and recommendation models plus inference for its apps — to beat a general-purpose GPU on cost-per-useful-work. You give up flexibility (Iris won't run every workload the way a Nvidia chip will) and you buy back margin. At Meta's scale, even a modest cost reduction across 14 gigawatts of compute is worth billions.

This is the same playbook we've now seen from three different angles: a startup choosing inference-specialized silicon in DeepSeek Is Building Its Own AI Chip, Microsoft swapping frontier labs out of its own products in Microsoft Is Quietly Swapping OpenAI and Anthropic Out of Excel and Outlook, and the broader ASIC wave in Can Custom AI Chips Dethrone Nvidia?. The through-line: Nvidia's biggest customers are all quietly building the tools to buy less of what Nvidia sells.

The chip is only half of it — Meta is also chasing the model race

The same week the Iris memo surfaced, Meta made a second aggressive move: it jumped deeper into the AI coding market with Muse Spark 1.1, which AI chief Alexandr Wang called Meta's "strongest model for agentic and coding work yet." The pricing is pointedly cheap — $1.25 per million input tokens, $4.25 per million output tokens, and $20 in free credits for every new API account — a deliberate undercut of Anthropic and OpenAI, which Wang described as "very aggressive and attractive."

Put the two moves together and the strategy snaps into focus. Meta is attacking its AI cost structure from both ends: cheaper compute (Iris) and a cheaper model it controls (Muse Spark). The pressure driving all of this is Wall Street's growing impatience — Mark Zuckerberg is being pushed to show a return on tens of billions in AI spending, and Meta has so far trailed OpenAI, Anthropic, and Google on popular models and applications. Owning the chip and the model is how Meta hopes to turn that spending into durable margin instead of a permanent tax.

Illustration of Meta's two-pronged AI cost strategy: a custom chip and an in-house model both driving down costs

## Can Iris actually dent Nvidia?

Here's the honest answer: not directly, and not soon — but that's the wrong way to measure it.

Iris will not out-perform Nvidia's flagship accelerators on raw, general-purpose capability, and Meta will keep buying Nvidia chips for the workloads where they win. What Iris can do is remove Meta's most predictable, highest-volume tasks — recommendation and inference — from Nvidia's meter. Every workload Iris absorbs is a GPU Meta doesn't rent. Multiply that across the four planned MTIA generations and the entire hyperscaler cohort doing the same thing, and the threat to Nvidia isn't a knockout blow — it's a slow erosion of its terminal market, the very thing a premium growth multiple is supposed to price in.

So the realistic 2027 picture isn't "Meta dumps Nvidia." It's "Meta caps how much of its 14 gigawatts Nvidia gets to fill." For a company whose valuation assumes near-total ownership of AI compute, that ceiling matters more than any single benchmark. Iris doesn't have to win the chip war. It just has to make Meta a smaller customer — and it's on track to start doing exactly that in September.

Frequently asked questions (FAQ)

What is Meta's Iris chip? Iris is the codename for an in-house AI accelerator under Meta's MTIA (Meta Training and Inference Accelerator) program, moving into production in September 2026. It's built to run Meta's own workloads — ranking, recommendation, and inference — more cheaply than general-purpose GPUs.

Is Meta dropping Nvidia? No. Meta says it "still expects to spend plenty" on Nvidia and AMD GPUs. Iris is a hedge to reduce dependence and cost, not a full replacement.

Who actually builds the chip? Meta designs it in-house with Broadcom's help; TSMC manufactures it. Samsung supplies RAM, SanDisk storage, and Sumitomo Electric fiber-optic gear.

How much is Meta spending on AI? Its 2026 capital expenditure is guided at roughly $125–145 billion, with plans to roughly double computing capacity to about 14 gigawatts in 2027.

What's Muse Spark 1.1 and how does it fit in? It's Meta's updated coding/agentic AI model, priced aggressively ($1.25/M input, $4.25/M output tokens, $20 free credits) to undercut Anthropic and OpenAI — the model-side complement to Iris's compute-side cost cutting.

Key takeaways

  • Meta's in-house Iris chip (MTIA program) enters production in September 2026, targeting a doubling of compute to ~14 gigawatts by 2027.
  • A test chip validated in ~6 weeks; Broadcom designs, TSMC manufactures, with Samsung, SanDisk, and Sumitomo supplying.
  • The point is cost, not raw performance — trimming a GPU bill inside a $125–145 billion capex budget.
  • Meta paired it with Muse Spark 1.1, an aggressively priced coding model, attacking AI costs on both compute and model fronts.
  • Iris won't beat Nvidia head-on, but it can shrink Meta as a customer — eroding the terminal market Nvidia's premium valuation depends on.

How this was written Research and a first draft came together with AI's help; verification and the final pass were entirely human.


References