Meta AI 论文提出 MIRA,用外层元推理器与执行器拆分改进研究智能体的下一步决策
Elvis Saravia 介绍 Meta AI 关于研究智能体如何决定下一步研究方向的论文 MIRA。该方法将智能体拆为外层元推理器与新执行器:外层读取持续研究记录并写下一份工作指令,执行器负责执行,决策只发生在工作指令边界,作者在这些点训练 critic 预测剩余收益,再训练单一 actor-critic(MIRA-AC)同时评估部分进展并选择下一步研究。
Banger paper from Meta AI on research agents that decide what to investigate next.
(bookmark it)
If you run long-horizon research agents, choosing the next investigation is hard to learn, because those decisions are rare in long traces and their effects show up several steps later.
MIRA splits the agent into two.
An outer meta-reasoner reads a persistent research record and writes a work order for the next investigation.
A fresh executor carries out each work order.
Decisions only happen at work-order boundaries, so the authors train a critic at those points to forecast remaining return, then a single actor-critic (MIRA-AC) that both values partial progress and picks the next investigation.
Even without training, the split improves theorem proving and open-ended architecture research.
Trained on the model's own proxy signals, MIRA-AC improves gold scores in all four autoresearch environments.
Paper: arxiv.org/abs/2610.02525
Chat with Paper: academy.dair.ai/papers/learn…
来源:dair_ai · x.com