跳到正文
原文
dair_ai· @dair_ai · X·· 21 小时前AI 评分57

NVIDIA 论文:NeMo-DCR 将 1T 模型 RL 训练权重同步从 87.5 分钟压缩到 150 秒

AI 导读

DAIR.AI 介绍 NVIDIA 的一篇论文:在所测六个模型中,每次 RL 步骤后仅 0.6% 到 1.2% 的模型权重发生变化,但标准 refit 每次更新后仍会把完整 checkpoint 复制到 rollout 集群,1T 模型跨两个 AWS region 复制需 87.5 分钟。

正文 · 原文

Another great paper from NVIDIA.

They find that only 0.6% to 1.2% of model weights change after each RL step in the six models they measured.

Yet a standard refit copies the full checkpoint to the rollout cluster after every update. For a 1T model across two AWS regions, that copy takes 87.5 minutes.

NVIDIA's NeMo-DCR sends only the changed values and still gives the rollout cluster the exact same bits as a full copy.

It maps changes from training shards straight into the checkpoint layout, encodes them as XOR masks or overwrites, and streams them through a relay tree while the delta is still being built. A joint commit and retries handle failures in the middle of a transfer.

A 1T refit at a 3% change rate takes 150 seconds instead of 87.5 minutes. Across 30B to 1T models, it is 12x to 40x faster than full-checkpoint transfer.

Paper: academy.dair.ai/papers/nemo-…

来源:dair_ai · x.com