跳到正文
原文
SemiAnalysis_· @SemiAnalysis_ · X·· 2 天前AI 评分38

SemiAnalysis 解释 LLM 自回归生成与输出成本

AI 导读

SemiAnalysis 解释 LLM 自回归生成机制:模型逐 token 预测下一个词,最后一层实为词表上的分类器,每步输出概率分布并采样。输出 token 价格是输入的 3-5 倍,因输入可并行处理,而输出需逐个生成,每步都要重读全部模型权重和 KV cache,属内存瓶颈。

正文 · 原文

An LLM generates text autoregressively. It predicts the next token, then repeats:

🟠 <start> -> "The" -> "The capital" -> "The capital of" -> "The capital of New" -> "The capital of New York" -> "The capital of New York is" -> "The capital of New York is Albany" -> "The capital of New York is Albany."

1 word = 1 token for simplicity sake but actual tokenizers may split words into multiple sub-word pieces.  The last layer of an LLM is technically a classifier over the vocabulary set but it runs that classification step over and over again.  At each step it outputs a probability distribution over every token and samples from it.

This is why output tokens are always 3-5x more expensive than input ones.  Input tokens are processed in parallel.  Output tokens are generated one at a time during decode, which is memory bound because it must re-read all the model weights and KV cache at each step. (2/3)

来源:SemiAnalysis_ · x.com