跳到正文
原文
rohanpaul_ai· @rohanpaul_ai · X·· 2 天前AI 评分45

清华 TokenRouter 实现逐 Token 模型路由,吞吐最高提升 64.15 倍

AI 导读

清华大学论文提出 TokenRouter,一个支持逐 Token 大小模型路由的 LLM 服务系统,在 5 种路由方法下吞吐量较现有最强方案提升 2.01 至 64.15 倍。针对 vLLM、SGLang 等框架一请求一模型导致的慢模型拖累问题,TokenRouter 为每个模型分配独立服务器,通过传递半成品回答并同步 KV cache,同时短暂缓存请求以增大各模型的批处理规模。

正文 · 原文

New Tsinghua paper builds TokenRouter, a serving system that runs per-token small-and-large model routing at upto 64.15X the throughput of existing setups.

Current popular serving frameworks (like vLLM and SGLang) run one model per request, so when two models share an answer, every step waits for the slower one.

TokenRouter gives each model its own server and lets them pass work back and forth. So, it hands a half-written answer between them while keeping the model’s memory of the text so far (the KV cache), and holds requests for a moment so each model works on bigger batches.

Across 5 routing methods, throughput rose 2.01 to 64.15 times over the stronger existing setup.

– arxiv. org/abs/2610.12242

Title: "TokenRouter: Efficient Serving System for Token-Level LLM Routing"

来源:rohanpaul_ai · x.com