微软论文:无中心管理者的编码智能体团队规模越大得分越高
微软一篇新论文提出,编码智能体无需中心管理者:在 ProgramBench 的 200 项任务中,让智能体自行认领任务的无管理者团队规模从 1 扩展到 128 个智能体时,平均得分在每一步都有提升,即使在最难的 5 项任务上也是如此。
编辑精选:介绍通过共享工具协调的无中心编码智能体团队,具有明确的工程实现与评测场景。论文转述尚未提供大团队成本,不能仅凭得分判断性价比。
Turns out coding agents don't need a boss:
New Microsoft paper finds that bigger teams of coding agents score higher and get there sooner when agents claim their own tasks without a central manager.
Even on the 5 hardest of ProgramBench's 200 tasks, scaling a manager-free team from 1 to 128 agents raised the average score at every step.
so for big jobs, add agents and let them coordinate through shared tools with no lead agent.
Popular multi-agent coding tools send all work through 1 lead agent, which can only manage so many helpers.
Agensh drops the lead agent. Each agent claims a sub-task, builds and tests it, merges it into a shared Git repo, and logs findings on a shared board.
Every run used 1 model, and the paper does not report what large teams cost.
来源:rohanpaul_ai · x.com