NVIDIA Revives "Rubin CPX" AI Accelerator with HBM4 Memory for Advanced AI Workloads

NVIDIA has reactivated its "Rubin CPX" AI accelerator project, introducing significant upgrades to its memory architecture. According to respected supply chain analyst Ming Chi Kuo, the revived Rubin CPX will now feature HBM4 memory, replacing the previously planned GDDR7. This strategic shift is designed to address the growing demands of large-scale agentic AI workloads, particularly those involving large language models (LLMs) and complex context management.

Enhanced Memory Architecture for Next-Generation AI

The original Rubin CPX design included 128 GB of GDDR7 memory. In its updated form, the accelerator now boasts 168 GB of HBM4 memory, offering higher bandwidth and faster data processing capabilities. This upgrade is crucial for handling the intensive prefill workloads required by modern LLMs, which rely on rapid context input and key-value (KV) cache processing.

The Rubin CPX GPUs are deployed in dedicated racks, with configurations ranging from 64 to 256 standalone units per rack. Each rack tray, equipped with eight CPX GPUs, can manage up to 1.34 TB of long-context prefill and KV cache, leveraging the speed and efficiency of HBM4 memory. This architecture enables faster and more efficient handling of large-scale AI tasks.

Separation of Prefill and Decode Tasks

NVIDIA’s new approach separates prefill and decode operations to optimize performance. The CPX GPUs are dedicated to prefill workloads, managing the input context and KV cache for LLMs. Meanwhile, standard Rubin GPUs are responsible for decode tasks. This division allows for specialized processing, resulting in faster token delivery and improved throughput for AI models with trillions of parameters.

To maximize efficiency, NVIDIA recommends maintaining a 1:1 ratio of CPX GPUs to regular GPUs in the decode racks. This balanced configuration ensures that both prefill and decode processes are adequately resourced, supporting seamless AI inference at scale.

Advanced Packaging and Compute Performance

The transition to HBM4 required a redesign of the Rubin CPX package. NVIDIA is likely utilizing TSMC’s advanced CoWoS-S or CoWoS-L packaging technologies to integrate the high-bandwidth memory. The monolithic die delivers an impressive 30 PetaFLOPS of NVFP4 compute performance, making it one of the most powerful AI accelerators available.

In addition to its compute capabilities, the Rubin CPX includes four integrated NVENC and four NVDEC video encoders on-chip. This integration streamlines multimedia workflows, eliminating the need for external video processing and further enhancing the accelerator’s versatility for AI-driven applications.

Implications for Large-Scale AI Deployments

As AI models continue to grow in complexity and scale, the need for specialized hardware becomes increasingly important. NVIDIA’s decision to revive and upgrade the Rubin CPX project demonstrates its commitment to meeting the evolving requirements of enterprise AI. By optimizing memory architecture and separating key processing tasks, the Rubin CPX is poised to deliver significant performance gains for organizations deploying advanced LLMs and agentic AI systems.