Meta accelerates timeline for its new generation of custom AI inference silicon

Meta has revealed an ambitious plan to expand its custom AI silicon portfolio, with a four-generation roadmap that aims to power large-scale inference workloads across its data centres, signalling a major shift in its AI hardware strategy.

Meta used its appearance at Hot Chips 2026 to set out a more ambitious timetable for its custom AI silicon, with the company now treating inference hardware as a regular product line rather than a side project. According to Meta’s March roadmap, the MTIA family will span four generations , MTIA 300, 400, 450 and 500 , with the chips intended to support ranking, recommendations and broader generative AI workloads across the company’s data centres. The scale of the effort reflects Meta’s need to run large volumes of inference efficiently, particularly for systems that must answer millions of requests each day.

The company’s earlier MTIA 300 chip, which is already in production, was built around recommendation-model training and was designed to cope with the memory-heavy nature of that workload. Meta said the chip combines purpose-built processing with large HBM3e capacity and high cache bandwidth, helping it match GPU performance at a more competitive total cost of ownership. Industry reporting also indicates that Meta has been working with Broadcom on the programme and has adopted a rapid release cadence, with new generations planned at roughly six-month intervals.

MTIA 400 is the clearest sign that Meta is widening the scope of the project. At Hot Chips, the company described the chip as an inference accelerator for more general use, not only recommendations. It uses multiple chiplet types, including compute, SoC and I/O chiplets, and adds hardware support for low-precision formats such as MXFP4. Meta also highlighted a larger memory system, including eight stacks of HBM3e and two forms of embedding cache to keep frequently used data close to the accelerator. The package is designed for scale, with a single tray holding four accelerators and a scale-up domain that can reach 72 ASICs.

On the silicon itself, Meta said MTIA 400 has an expanded compute core with an 8×6 grid of processing elements, a separate special function unit for non-GEMM operations, and RISC-V-based control cores inside each processing element. The uncore adds a 2D mesh interconnect, congestion control, Ethernet-based scale-up fabric and PCIe-based scale-out networking. Meta says the chip delivers 12 PFLOPS of FP4 performance, while compared with MTIA 200 it offers more than 15 times the FP16 compute and 46 times the DRAM bandwidth. Successors are already in development: MTIA 450 is aimed at stronger GenAI inference performance, while MTIA 500 is intended to push scale further, in step with the industry-wide move towards ever larger accelerator domains. Meta’s broader strategy appears to be clear enough: use custom silicon where it fits best, keep buying external accelerators where necessary, and build its own inference stack around the workloads it runs most often.

Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.