Apple 论文:强编码智能体自主做 ML 工程,最小 harness 表现最好
Rohan Paul 转述 Apple 新论文,指出自动 ML 工程的进展来自模型与运行时而非外围脚手架。论文用同一代码库、模型、硬件和 24 小时预算重跑各种封装,发现一个带 shell 和文件访问的单一编码智能体匹配或超过 4 个多智能体 ML 系统,给模型 shell 而非聊天框是唯一明显起作用的改动。
New Apple paper basically says progress in automated ML engineering has come from models and runtimes rather than the scaffolding built around them.
It finds one well-prompted coding agent with shell and file access matched or beat 4 multi-agent ML systems, so skip the orchestration and start with a single session.
Those systems add search trees, memory layers, and specialist agent teams.
They were designed when a model could only write code, not run it.
This paper re-ran every wrapper on 1 codebase with the same model, hardware, and 24-hour budget. Giving the model a shell instead of a chat box was the only change that clearly mattered.
With GLM 5.2, the minimal agent medaled on 62.5% of Kaggle tasks against 47.1% for the best published harness. Adding parallel agents and a message channel actually dropped its medal rate from 55.7% to 33.3%.
– arxiv. org/abs/2609.40303
Title: "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?"
来源:rohanpaul_ai · x.com