SemiAnalysis 解释 LLM 自回归生成与输出成本
SemiAnalysis 解释 LLM 自回归生成机制:模型逐 token 预测下一个词,最后一层实为词表上的分类器,每步输出概率分布并采样。输出 token 价格是输入的 3-5 倍,因输入可并行处理,而输出需逐个生成,每步都要重读全部模型权重和 KV cache,属内存瓶颈。
An LLM generates text autoregressively. It predicts the next token, then repeats:
🟠 <start> -> "The" -> "The capital" -> "The capital of" -> "The capital of New" -> "The capital of New York" -> "The capital of New York is" -> "The capital of New York is Albany" -> "The capital of New York is Albany."
1 word = 1 token for simplicity sake but actual tokenizers may split words into multiple sub-word pieces. The last layer of an LLM is technically a classifier over the vocabulary set but it runs that classification step over and over again. At each step it outputs a probability distribution over every token and samples from it.
This is why output tokens are always 3-5x more expensive than input ones. Input tokens are processed in parallel. Output tokens are generated one at a time during decode, which is memory bound because it must re-read all the model weights and KV cache at each step. (2/3)
来源:SemiAnalysis_ · x.com