跳到正文
原文
qvac· @qvac · X·· 2 天前AI 评分58

llama.cpp 合并面向 Apple Silicon 投机解码的新 Metal kernels

AI 导读

qvac 宣布其提交的 PR 已被合并进 llama.cpp,为 Apple Silicon 投机解码带来新的 Metal kernels。此前在 Mac 上投机解码比普通解码更慢,此次改动后在 M3 Ultra 上最高可达 3.4 倍加速,达到 110 tok/s 对比 32.1 tok/s。作者感谢 @ggerganov 的评审与优化。

正文 · 原文

One of our PR just got merged into llama.cpp: new Metal kernels for speculative decoding on Apple Silicon.

Before this PR, speculative decoding on a Mac was slower than plain decoding. On an M3 Ultra it now runs up to 3.4x faster than plain decoding, 110 tok/s against 32.1.

Thanks to @ggerganov for reviewing and refining it.

github.com/ggml-org/llama.cp…

来源:qvac · x.com