跳到正文
原文
mingchikuo· @mingchikuo · X·· 2026-08-31AI 评分66

郭明錤:Nvidia 恢复 Rubin CPX 项目,预计 1Q27 量产

AI 导读

郭明錤表示,行业调研显示 Nvidia 已恢复此前被认为被砍掉的 Rubin CPX 项目,预计 1Q27 量产,且 prefill 性能较旧设计更强。

正文 · 原文

Just as the market had come to believe that Rubin CPX had been dropped from Nvidia’s product roadmap, my latest industry checks indicate that Nvidia has revived the program, with production expected to begin in 1Q27. Compared with the previous design, the revived Rubin CPX delivers stronger prefill performance and features major changes to both its GPU specifications and rack architecture, underscoring the high priority Nvidia places on prefill solutions.

Key changes:

1. Rubin CPX GPU specs

CPX delivers near-Rubin compute performance and matches Rubin’s maximum power rating of 2,300 W per GPU. CPX moves to 168 GB of HBM4, vs. 288 GB on Rubin and 128 GB of GDDR7 on the previous CPX design.

2. Rack design

The new CPX uses a standalone MGX ETL rack rather than sharing a rack with Rubin, as in the previous design. Customers can opt for 64, 128, 192, or 256 CPX GPUs depending on their needs. Within a CPX rack, each group of 64 CPX GPUs forms a rack module comprising eight compute trays (eight CPX GPUs per tray) and one switch tray.

3. Scale-up and scale-out

NVLink is used only for scale-up among the eight CPX GPUs within each tray, with 1–1.5 TB/s of NVLink bandwidth per CPX (vs. 3.6 TB/s per Rubin). Inter-tray scale-out within each rack module runs over Spectrum-6 Ethernet using all-copper L1 links. Across rack modules, scale-out is handled by each module’s Spectrum-6 switch over OSFP optical links.

4. How it works

CPX must be paired with Vera Rubin NVL72, and Nvidia recommends a 1:1 ratio of CPX to Rubin GPUs. CPX handles prefill and builds the KV cache, which is then transferred to Rubin over Ethernet RDMA for decode.

5. Product positioning: Best performance per dollar for long-context prefill

Over 50% of today’s AI inference workload comes from processing input context and building the corresponding KV cache. CPX therefore offers a more flexible, lower-cost way to handle prefill. Each eight-CPX tray has approximately 1.34 TB of HBM4, sufficient for most long-context prefill workloads and associated KV cache requirements.

来源:mingchikuo · x.com