跳到正文

#论文/研究

今日 6 条
今天10月11日周日
  1. rohanpaul_ai56

    Meta 新论文提出 GitSwarm,让多个相同智能体在共享 Git 仓库上协作,无中心分配任务,每个智能体阅读过往工作、提交结果并标注所依赖的先前提交(含失败尝试)。在 ProgramBench 的 50 项程序重建任务上,GitSwarm 得分 79.4%,而同等算力下被告知持续工作的单个 Codex 智能体最高为 65.1%;后续智能体在 94.7% 的已保存贡献基础上继续构建。

10月10日周六
  1. rohanpaul_ai45

    清华大学论文提出 TokenRouter,一个支持逐 Token 大小模型路由的 LLM 服务系统,在 5 种路由方法下吞吐量较现有最强方案提升 2.01 至 64.15 倍。针对 vLLM、SGLang 等框架一请求一模型导致的慢模型拖累问题,TokenRouter 为每个模型分配独立服务器,通过传递半成品回答并同步 KV cache,同时短暂缓存请求以增大各模型的批处理规模。

  2. rohanpaul_ai51

    New NYU and Amazon paper finds that distilling a few skills that keep producing a useful training signal matches or beats distilling a skill bank up to 11× larger. Skills are short written tips, like a rule for counting cases, that a model absorbs by learning from a copy of itself that reads them. Picked by topic match, under 25% of them gave any useful signal across 3 Qwen models. SGUID keeps only skills that help early in training and still help late. With 6 such skills, 3 of 4 models matched or beat the full bank of 30 to 71 skills on math contest tests. A 2nd round with 3 new skills lifted Qwen3-8B from 64.3% to 66.3%.

10月9日周五
  1. thexpin52

    Seed 团队一篇预印本报告 DeepSeek V4 和 V4.1-Flash 存在相位敏感性:相同信息在压缩 KV-cache 块中的位置不同,检索难度也不同。该技术可降低内存和注意力开销,但长上下文检索准确率在不同位置间差异最多达 40 个百分点;作者指出平均分会掩盖这类反复出现的弱点,且该发现针对检索问题,并非解释所有已报告的模型异常。论文见 arxiv.org/abs/2609.36322。