跳到正文
原文
LysandreJik· @LysandreJik · X·· 2 天前AI 评分59

TRL v1.15 发布:默认启用 fused LM head 大幅降低后训练显存

AI 导读

TRL v1.15 发布,SFT、DPO、KTO、GRPO、RLOO 和蒸馏默认改用 fused LM head,用 Triton kernel 直接计算所需 token 级量,不再生成巨大的 [batch, seq, vocab] logits 张量。

正文 · 原文

TRL v1.15 is out, and it’s an absolute banger of a release for memory-efficient post-training.

The main change: SFT, DPO, KTO, GRPO, RLOO and Distillation now use a fused LM head by default.

Instead of materializing the huge [batch, seq, vocab] logits tensor, a Triton kernel computes the token-level quantities we actually need directly.

The results are significant 👇

On Gemma 3 1B with a 262k vocabulary, on the same GPU:

DPO: 10k → 59k max sequence length

KTO: 9k → 63k

GRPO: 28k → 114k

RLOO: 23k → 100k

SFT: 20k → 107k

Up to 6.9x longer sequences, with peak memory at 8k context reduced by 52-82%.

Speed is not negatively impacted: training is up to ~11% faster.

Nothing to enable: this is now the default in TRL v1.15!

There’s more in the release too: selective activation checkpointing for SFT, assistant-only loss for vision datasets, better conversation logging, improvements to AsyncGRPO / AsyncDistillation, and a long list of fixes.

I really like optimizations like this: the training API doesn’t need to become more complicated as the implementation underneath gets much better.

github.com/huggingface/trl/r…

来源:LysandreJik · x.com