AgentWorld 论文:多智能体团队不到三分之一的行动真正有助于完成任务
More agents don't mean higher performance. There is a coordination bottleneck to consider. Not to mention the unnecessary costs. So how many of a multi-agent team's actions actually help it finish the task? In this AgentWorld paper, fewer than a third. AgentWorld puts 3 to 20 LLM agents with different roles into a game sandbox for tasks that run 50+ rounds. Agents can't see each other's internal state, so they have to coordinate through messages and shared plans. Gemini 3 Flash has the highest task success at 52.0%. Coordination tasks are the hardest category, at 12% success, and common failures include communication breakdowns, role confusion, and lost shared plans. Paper: https://arxiv.org/abs/2609.31590 Chat with Paper: https://academy.dair.ai/papers/agentworld-benchmarking-long-horizon-collaboration-of-multi-agent-llms-2609.31590
AgentWorld 将 3 到 20 个不同角色的 LLM 智能体放入游戏沙盒,执行 50+ 轮的长程任务,结果显示多智能体团队中不到三分之一的行动真正有助于完成任务。Gemini 3 Flash 任务成功率最高,为 52.0%;协调类任务最难,成功率仅 12%,常见失败包括沟通中断、角色混淆和共享计划丢失。
来源:X:Elvis Saravia (@omarsar0, DAIR.AI) · x.lingyaoai.com