清华论文评测AI智能体从对局回放学习游戏策略:可登顶人类榜单但复杂规则下多停滞
清华一篇论文构建了 AAArena,包含来自清华年度 bot 构建比赛的 12 个游戏和 1,920 个存档人类程序作为对手,让一个编码智能体在模型权重不变的前提下读规则、选对手、研究回放并在比赛预算内重写自己的 bot。
New Tsinghua paper finds that AI agents improving game bots from match replays can top human leaderboards, but mostly stall on games with complex rules.
Getting AI to learn a winning game strategy from a limited number of matches is still hard, especially against changing rivals.
They built AAArena from 12 games in Tsinghua's yearly bot-building contest, with 1,920 archived human programs as rivals. A coding agent, with its model weights unchanged, reads the rules, picks opponents, studies replays, and rewrites its bot within a match budget.
Detailed replays beat win/loss-only feedback in all 3 games tested. With replays, a Pacman bot reached rank 1, versus rank 11 without them.
Tripling the match budget did not push any of 4 stuck bots to rank 1.
– arxiv. org/abs/2610.12341
Title: "Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition"
来源:rohanpaul_ai · x.com