跳到正文
原文
X:Rohan Paul (@rohanpaul_ai)· @rohanpaul_ai·· 4 小时前AI 评分60

美中实验室论文提出 AgentBug-Smith,将智能体框架 bug 转为可运行测试

Self-improving AI agents will need to fix their own code. And this paper from top US+China labs, shows coding agents miss most such bugs but improve with lessons from past fixes. that real bugs in agent harnesses, can be automatically turned into a growing set of runnable tests. An agent's own code is everything around the model: tool calls, memory, and prompts. Its bugs depend on live model calls, which makes them hard to recreate and test. So the researchers built AgentBug-Smith, which turns real GitHub bug reports into runnable tests. The result is a 200-bug benchmark that keeps growing. The best of 3 coding agents fixed just 9% of those bugs, versus about 40% reported on regular software bugs. A short guide of lessons from past fixes lifted an agent from 1 to 6 correct fixes on 79 unseen bugs. Before trusting a coding agent with your agent's code, try it on bugs you've already fixed.

AI 导读

美中多家实验室的论文提出 AgentBug-Smith,把 GitHub 上智能体框架的真实 bug 报告自动转成可运行测试,形成一套 200 个 bug 且持续增长的基准。

来源:X:Rohan Paul (@rohanpaul_ai) · x.lingyaoai.com