Qualcomm goes all-in on inferencing with purpose-built cards and racks

News
Oct 27, 20256 mins

Qualcomm’s AI200 and AI250 move beyond GPU-style training hardware to optimize for inference workloads, offering 10X higher memory bandwidth and reduced energy use.

AI data center power use
Credit: Shutterstock/Gorodenkoff

It’s becoming increasingly clear that AI inferencing — where a trained model uses new, unseen data to make predictions — demands a fundamentally different hardware profile than AI training.

Where training is compute-intensive and time-tolerant, inference is latency-sensitive, memory-bound, and continuous.

Recognizing this, semiconductor giant Qualcomm has announced a new generation of accelerator cards and racks purpose-built to handle inference at scale. The new AI200 and AI250 feature near-memory computing, high memory capacity, and, it says, high per-dollar/per-watt and rack-scale performance.

The move underscores Qualcomm’s growing seriousness about AI infrastructure; it is also indicative of a broader industry pivot to purpose-built platforms, where technology is customized to specific workloads “right from the silicon, to the software, to the developer environment,” explained Brian Jackson, a principal research director at Info-Tech Research Group.

Next-gen memory architecture, lower TCO

According to Qualcomm, the new AI200 and AI250 architectures offer “rack-scale performance” and “superior memory capacity” for data center AI inferencing. AI250 introduces what the company calls an “innovative memory architecture,” based on near-memory computing (NMC). As its name implies, NMC places the processor chips as close to, or within, a memory system to minimize the need for data transfer and allow it to be processed where it resides.

This capability in Qualcomm’s architecture provides 10X higher effective memory bandwidth with significantly lower power consumption, Qualcomm says. This supports disaggregated AI inferencing (in which inferencing is split into stages) to make more efficient use of hardware and improve performance.

The Qualcomm AI200 is purpose-built to deliver lower total cost of ownership (TCO) and better performance for large language models (LLM), large multimodal model (LMM), and generative AI workload inferencing. The system supports 768 GB of low-power double data rate (LPDDR) memory per card for higher memory capacity and lowers costs, according to the company. And both AI200 and AI250 are compatible with leading AI frameworks and feature scale up and scale out capabilities, support for confidential computing, direct liquid cooling for thermal efficiency, and have low rack-level power consumption of 160 kW.

Developers can deploy Hugging Face models with a single click and easily access Qualcomm’s transformer libraries, APIs, inferencing suite, ready-to-use AI apps and agents, and other tools.

Qualcomm said it is committed to a multi-generation data center roadmap with an “annual cadence,” focusing on performance, energy efficiency, and TCO.

This signals its intent to keep pace with the market and “demonstrates Qualcomm’s seriousness about playing in this space,” noted Matt Kimball, VP and principal analyst at Moor Insights and Strategy.

Early partnership with Saudi Arabia’s Humain

The new Qualcomm AI200 and AI250 cards and racks are expected to be commercially available in 2026 and 2027, respectively.

However, Qualcomm has already lined up its first customer: Saudi Arabia-based Humain, whose models will be integrated with Qualcomm’s chip and rack tools to target 200 megawatts for high-performance inference services. The companies say it will position Humain as the world’s first fully optimized edge-to-cloud hybrid AI.

“By focusing on AI inferencing rather than training, Humain is targeting a fast-growing need in the AI space, with demand coming from enterprises of all shapes and sizes,” said Info-Tech’s Jackson.

Performance and security are priorities in inferencing; users want a fast and accurate response that can’t be tampered with by outside actors, he pointed out.

“Supported by Qualcomm’s new accelerator hardware, Humain will be looking to host enterprise client applications with heavy AI workloads, delivering them in a cloud model,” said Jackson.

New hardware needed for inferencing

Scott Young, a principal advisory director at Info-Tech Research Group, pointed out that training is compute-bound and time-tolerant, whereas inference is memory and/or latency-bound and “presumably continuous.”

These require different types of hardware, he noted. Training hardware is expensive, and although it can be used for inference, it may not always be optimal.

Demand for inference will continue to grow as enterprises deploy more agentic AI, he said. “Hardware specifically targeting inference can be more cost-effective, and can be appropriately scaled for rapid growth in usage of these AI models,” Young explained.

From a strategy perspective, there is a longer term enterprise play here, noted Moor’s Kimball; Humain is Qualcomm’s first customer, and a cloud service provider (CSP) or hyperscaler will likely be customer number two. However, at some point, these rack-scale systems will find their way into the enterprise.

“If I were the AI200 product marketing lead, I would be thinking about how I demonstrate this as a viable platform for those enterprise workloads that will be getting ‘agentified’ over the next several years,” said Kimball.

It seems a natural step, as Qualcomm saw success with its AI100 accelerator, a strong inference chip, he noted. Right now, Nvidia and AMD dominate the training market, with CUDA and ROCm enjoying a “stickiness” with customers.

“If I am a semiconductor giant like Qualcomm that is so good at understanding the performance-power balance, this inference market makes perfect sense to really lean in on,” said Kimball.

He also pointed to the company’s plans to re-enter the datacenter CPU space with its Oryon CPU, which is featured in Snapdragon and loosely based on technology it acquired with its $1.4 billion Nuvia acquisition.

Ultimately, Qualcomm’s move demonstrates how wide open the inference market is, said Kimball. The company, he noted, has been very good at choosing target markets and has seen success when entering those markets. “That the company would decide to go more ‘in’ on the inference market makes sense,” said Kimball. He added that, from an ROI perspective, inferencing will “dwarf” training in terms of volume and dollars.