Arena CEO:AI 模型对话超过 20 轮后失准率至少 50%
Arena CEO @ml_angelopoulos 表示,AI 模型在对话超过 20 轮后,出现失准(misaligned)的比例至少为 50%。他称其评测代表模型部署后在真实用户手中的安全与对齐状况,能获得红队和基准测试拿不到的数据;相关曲线均呈凸且上升趋势,对话越长越可能至少出现一次欺骗、越权操作或错误归因。他指出,在真实自然的用户任务分布中,目前这些模型尚未表现出对齐。
借 Arena CEO 之口给出对话轮次与失准率的量化关系,可帮助读者理解部署后安全评测与红队的差异。
Arena CEO @ml_angelopoulos reveals that AI models show misalignment in at least 50% of conversations once they pass 20 turns:
"The thing that makes our evals different is that they represent the real world post-deployment safety and alignment of these AI models."
"They represent what's actually happening when you put them in users' hands, and because of that, we're able to get all sorts of interesting data that you wouldn't get through red teaming and that you wouldn't get from a benchmark."
"All these curves are convex and increasing, which means that the longer you go in conversation, the more likely it is that you're going to be witnessing at least one deception, at least one unauthorized action, at least one false attribution."
"Once you start getting into a 20-plus turn conversation, the rate at which models are misaligned is at least 50%."
"There's still so far to go to make sure that in the distribution of actual organic user tasks, these models exhibit alignment because today they do not."
@arena
来源:MTSlive · x.com