(Source: Cerebras Systems Inc.)
  • AMD (NASDAQ:AMD) and Cerebras (NASDAQ:CBRS) are collaborating to advance a workload-optimized approach to ultra-low-latency AI inference infrastructure
  • AMD Helios and the Cerebras Wafer-Scale Engine will operate as a single disaggregated inference workflow, combining throughput from AMD Instinct GPUs with fast token generation of Cerebras Wafer-Scale Engine
  • Cerebras plans to deploy AMD Helios in its data centres, with the joint solution expected to be available first through Cerebras Cloud in the second half of 2026
  • Advanced Micro Devices stock (NASDAQ:AMD) last traded at US$554.42, and Cerebras Systems stock (NASDAQ:CBRS) last traded at US$221.60

Advanced Micro Devices (NASDAQ:AMD) and Cerebras Systems (NASDAQ:CBRS) have announced a technical partnership to develop a new disaggregated artificial intelligence inference solution that combines AMD’s Helios rackscale infrastructure with Cerebras’ Wafer-Scale Engine technology.

The companies unveiled the joint offering at Advancing AI 2026, positioning it as a platform designed to deliver ultra-low latency AI inference while improving throughput and energy efficiency for advanced AI workloads.

According to AMD and Cerebras, the integrated solution will combine AMD Helios systems as a high-throughput processing engine with Cerebras’ Wafer-Scale Engine technology for low-latency token generation and decoding. The companies said the architecture is expected to deliver up to five times higher tokens per second per watt (T/s/W) compared with conventional approaches.

This article is a journalistic opinion piece that has been written based on independent research. It is intended to inform investors and should not be taken as a recommendation or financial advice.

“AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach,” Dr. Lisa Su, AMD’s chair and CEO, said in a news release. “AMD Helios delivers leadership performance and scale for the broadest range of inference workloads. Together with Cerebras, we are extending that leadership into the most latency-sensitive applications and creating a powerful new platform for real-time agentic AI.”

The announcement comes as AI inference workloads become increasingly diverse, requiring different combinations of latency, throughput, scalability and cost efficiency. While some enterprise and cloud deployments focus on maximizing overall token generation, other applications such as coding assistants, real-time copilots, live AI agents and autonomous workflows place a premium on response speed.

To address these varying requirements, AMD and Cerebras are adopting a disaggregated inference approach, in which different stages of the AI inference process are optimized independently using specialized compute architectures.

Under the proposed workflow, AMD Helios systems will handle prompt processing and large context windows, areas that typically require significant throughput and scalability. Cerebras’ Wafer-Scale Engine technology will be responsible for token generation and decoding, tasks that are heavily dependent on memory bandwidth and low-latency execution.

By connecting the two systems within a single inference workflow, the companies aim to provide customers with a platform that can deliver rapid response times without sacrificing throughput or deployment scale.

The partners said demand for faster token generation is expected to increase as AI expands into applications such as software development, autonomous agents, robotics and scientific research, where response times can directly affect system effectiveness and user experience.

AMD described Helios as providing the rack-scale efficiency and throughput needed to process large volumes of complex AI requests, while Cerebras highlighted its Wafer-Scale Engine technology’s ability to return generated tokens in real time.

The companies said the resulting platform is specifically designed for the ultra-low-latency segment of the AI inference market, while also supporting balanced inference workloads across modern data centre environments.

As part of the partnership, Cerebras Systems Inc. plans to deploy AMD Helios systems within its own data centres. The companies expect the joint solution to be made available initially through Cerebras Cloud during the second half of 2026.

The announcement reflects a broader industry trend toward heterogeneous AI infrastructure, where multiple specialized computing technologies are combined to address the growing complexity of inference workloads and the increasing performance demands of enterprise AI applications.

Meanwhile, this announcement also comes as AMD launched its next-generation AI infrastructure and physical AI portfolio at Advancing AI 2026, led by AMD Helios rackscale solutions, now in production to be deployed by leading AI companies at gigawatt scale.

Advanced Micro Devices stock (NASDAQ:AMD) closed trading 0.38 per cent higher at US$554.42 and has risen roughly 160 per cent since the year began.

Cerebras Systems stock (NASDAQ:CBRS) closed trading 5.62 per cent higher at US$221.60, but it is down more than 130 per cent since the year began.

Join the discussion: Find out what the Bullboards are saying about chip stocks like AMD and check out Stockhouse’s stock forums and message boards.


More From The Market Online

@ the Bell: Tesla plunges as markets turn lower

Canadian stocks moved lower on Thursday, mirroring losses in US markets as Brent crude oil prices...

StockTalk | Cannabis Report: In the pocket

Tilray Brands (TSX/NASDAQ:TLRY), one of the largest players in the cannabis industry, is capitalizing on the...
Teck employees at its Quebrada Blanca operation in Chile

Teck Q2 beats estimates as copper drives earnings

Teck (TSX:TECK) reported adjusted EBITDA of $2.2 billion in Q2 2026, more than tripling from a year ago, driven by higher copper production.