Nvidia vs. AMD vs. Cerebras: Which Is the Best AI Inference Stock to Buy Today?

1 month ago 8

Geoffrey Seiler, The Motley Fool

Thu, July 16, 2026 astatine 7:40 AM CDT 5 min read

While the archetypal signifier of the artificial quality (AI) inclination was each astir ample connection exemplary (LLM) training, the adjacent signifier is progressively becoming astir inference -- putting those trained models to enactment connected real-world tasks. This is expected to go the larger of the 2 markets eventually.

Chipmakers Nvidia (NASDAQ: NVDA), Advanced Micro Devices (NASDAQ: AMD), and Cerebras Systems (NASDAQ: CBRS) are each looking to go the person successful powering inference workloads. Excelling successful those computing tasks tends to beryllium much astir speedy entree to representation than earthy computing power, and each 3 are taking antithetic approaches to tackle this issue. As such, let's excavation into which looks similar the champion inference semiconductor banal to bargain close now.

Missed Nvidia successful 2009? This Rare Signal Is Flashing Again. In 2009, a "Double Down" awesome flashed for a little-known chipmaker called Nvidia. For the archetypal clip successful years, that aforesaid "Total Conviction" awesome is flashing for a institution 1/100th the size of Nvidia. Continue »

Nvidia

Nvidia has agelong been the AI infrastructure leader: Its graphics processing units (GPUs) person been the main chips utilized to bid AI models. The institution has developed a immense moat successful this arena by popularizing its CUDA bundle level with developers. Most foundational AI codification was written successful CUDA and optimized specifically for Nvidia's chips.

Inference is simply a full different ballgame, though, and the institution "acquired" Groq and its connection processing units (LPUs) and incorporated them into its CUDA ecosystem to beef up its inference offering. LPUs usage a tiny magnitude of on-chip SRAM (static random-access memory) to summation inference speeds. Nvidia has fundamentally created afloat server racks designed specifically for inference that usage a operation of LPUs and GPUs. GPUs packaged with high-bandwidth representation (HBM) instrumentality attraction of the prefill signifier of knowing a user's prompt, portion LPUs grip the decode signifier of giving a response. Since LPUs usage SRAM, they tin reply with astir zero lag.

Cerebras

Like Nvidia, Cerebras is besides utilizing on-chip SRAM to assistance tackle inference workloads and amended speeds. However, it is doing it successful a precise antithetic way. Because SRAM is physically bulky, each Nvidia LPU uses a tiny amount; to grip these workloads, galore LPUs request to beryllium interconnected successful a immense cluster. Cerebras, connected the different hand, has created elephantine wafer-sized chips astir the size of meal plates that enactment galore modular chips' worthy of hardware onto a azygous slab of silicon. The effect is simply a spot that is 6 times faster than Nvidia's LPUs and 15 times faster than its GPUs.

Read Entire Article