Founder Stack

Groq's LPU chips are forcing a referendum on wheth

Groq's LPU chips are forcing a referendum on whether inference needs Nvidia at all

· Semiconductors · Financial Times

Groq has spent nearly a decade building a case that runs against the grain of the current AI infrastructure boom: that inference, the process of actually running a trained model to answer a query, does not need to happen on the same general-purpose GPU architecture that dominates training. Founded by Jonathan Ross, who previously helped design Google's original Tensor Processing Unit, Groq built its Language Processing Unit, or LPU, as a chip purpose-built for the sequential, low-latency demands of running large language models rather than the parallel matrix-multiplication workloads GPUs were originally designed for. The pitch has found real commercial traction as inference costs, not training costs, have become the dominant line item for companies deploying AI products at scale. Groq's public benchmarks showing dramatically faster token generation speeds than GPU-based inference for popular open-weight models attracted partnerships with Saudi Arabia's Aramco Digital, which committed to building out a large LPU data center in the kingdom, and with Bell Canada and other telecom and sovereign-cloud customers looking to reduce dependence on Nvidia allocation. Those deals helped Groq raise a $640 million Series D in 2025 at a $2.8 billion valuation, and by early 2026 the company was reportedly in discussions for a substantially larger round that would value it well above that mark, reflecting investor appetite for any credible Nvidia alternative. The competitive and geopolitical context matters here as much as the technology. Export controls on advanced Nvidia chips to certain markets have made sovereign buyers, particularly in the Gulf, eager to diversify their AI infrastructure supply chains, and Groq's LPU, being a different chip architecture built by a US company, occupies an unusual middle position: attractive to governments wary of over-relying on any single vendor, but still subject to its own export-control questions given the strategic sensitivity of frontier compute. The harder problem for Groq is manufacturing and ecosystem lock-in. Nvidia's CUDA software stack and its scale advantages with foundry partner TSMC give it a moat that is only partly about raw chip performance; Groq has to convince developers to port workloads to a new toolchain and convince foundry partners to prioritize its wafer orders against far larger customers. Groq has managed this by focusing narrowly on inference rather than trying to compete across training and inference simultaneously, a scoping decision that has kept its capital needs lower than a full-stack competitor but also caps its addressable market to whatever share of AI spending inference represents. What to watch: whether Groq closes a mega-round in 2026 that funds a meaningful expansion of LPU manufacturing capacity, whether more sovereign AI buyers diversify away from pure-Nvidia stacks in ways that benefit Groq specifically, and whether Nvidia's own inference-optimized chip roadmap narrows the latency and cost gap Groq has built its business around.

Original source: Financial Times
Read more on Founder Stack