跳到正文
原文
rohanpaul_ai· @rohanpaul_ai · X·· 1 天前AI 评分66

Apple 论文:强编码智能体自主做 ML 工程,最小 harness 表现最好

AI 导读

Rohan Paul 转述 Apple 新论文,指出自动 ML 工程的进展来自模型与运行时而非外围脚手架。论文用同一代码库、模型、硬件和 24 小时预算重跑各种封装,发现一个带 shell 和文件访问的单一编码智能体匹配或超过 4 个多智能体 ML 系统,给模型 shell 而非聊天框是唯一明显起作用的改动。

正文 · 原文

New Apple paper basically says progress in automated ML engineering has come from models and runtimes rather than the scaffolding built around them.

It finds one well-prompted coding agent with shell and file access matched or beat 4 multi-agent ML systems, so skip the orchestration and start with a single session.

Those systems add search trees, memory layers, and specialist agent teams.

They were designed when a model could only write code, not run it.

This paper re-ran every wrapper on 1 codebase with the same model, hardware, and 24-hour budget. Giving the model a shell instead of a chat box was the only change that clearly mattered.

With GLM 5.2, the minimal agent medaled on 62.5% of Kaggle tasks against 47.1% for the best published harness. Adding parallel agents and a message channel actually dropped its medal rate from 55.7% to 33.3%.

– arxiv. org/abs/2609.40303

Title: "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?"

来源:rohanpaul_ai · x.com