跳到正文
原文
rohanpaul_ai· @rohanpaul_ai · X·· 1 天前AI 评分64

Stanford 论文提出 DeLM 去中心化多智能体框架,在编码任务上超越 Claude Code 和 Codex

AI 导读

Stanford 论文提出 DeLM,用共享上下文和任务队列替代中心协调智能体,让各智能体独立领取任务并即时共享发现。

正文 · 原文

A new Stanford paper just proved that decentralized multi-agent systems can beat Claude Code and Codex on complex coding tasks, running up to 2.49× faster!

DeLM tackles a key bottleneck in multi-agent systems: time wasted repeating work and waiting on other agents.

DeLM replaces the central coordinating agent with a shared context and task queue. Agents pick up tasks independently, share discoveries as they happen, and build on each other’s progress. When one agent finds a solution or hits a dead end, the others can use that information immediately.

DeLM beats state-of-the-art in both speed and accuracy on long-horizon tasks from Terminal-Bench 4.0, DeepSWE v1.1, and ProgramBench:

- Up to 2.49× faster execution than the vanilla Claude Code and Codex baselines

- Up to +19.2 percentage points in accuracy over the vanilla Claude Code and Codex baselines

- Up to +19.9 points in ProgramBench test pass rate within the same 120-minute budget

Built on Codex and Claude Code. Code and 720 trajectories are available so you can explore how the agents collaborate! An open-source plugin lets you try DeLM directly in Codex and Claude Code.

来源:rohanpaul_ai · x.com