Arena 测评 AI 智能体虚假归因行为
Arena 测评发现 AI 智能体存在"虚假归因"问题,即把用户未说过的话或未做的决定错误归给用户。GPT-6 Luna 和 Astra 的误引率最低(15.6% 和 28.6%),但错误归因率高达 53.1% 和 48.2%;GPT-6 Sol 则以 23.5% 的比例最常篡改用户历史记录。
We also looked into False Attribution, where the agent attributes a statement, request, choice, approval, or fact to the user, but user-provided evidence contradicts that attribution.
Interestingly we saw some models misquoting the user (misstating what a user asked for), while others credit the user with someone else's work.
There were varied patterns across models, with GPT-6 Luna and Astra rarely misquoting (15.6% and 28.6% respectively) but often misattributing (53.1% and 48.2%). Interestingly their sibling model GPT-6 Sol has the highest rate of misstaging the user’s history (23.5%).
来源:arena · x.com