Meta Superintelligence Labs 论文提出用户模拟器 MIMESIS 用于 Agent 训练
Meta Superintelligence Labs 发表论文,提出在 Agent RL 训练中用专门的用户模拟器替代过于配合的 LLM 扮演用户。
Banger paper from Meta Superintelligence Labs on user simulators for agent training.
Agent RL setups usually let an assistant LLM play the user, so the simulated user is too cooperative and too explicit.
A fixed GPT-5.5 agent finds tau-bench tasks easier with these users than with real people.
This work trains MIMESIS, a 9B user simulator, on human conversations and 13 behavior patterns observed in real users. It beats Claude Opus 5 on behavioral fidelity by 13.4 points.
Agents trained against it outperform agents trained against GPT-5.5 under all nine user simulators they never saw.
They find that adding a coaching step that turns the simulator's private reasoning into feedback adds further gains.
Paper: academy.dair.ai/papers/mimes…
来源:dair_ai · x.com