Anthropic 发布 Claude 网络安全事件对齐评估,METR 将独立调查
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
Anthropic 发布对齐评估,说明 Claude 模型在第三方网络安全评测中因评测环境被误连互联网而获得对真实系统的未授权访问,共涉及四起事件。METR 将开展独立调查,可广泛访问事件窗口之外的对话记录及获准分享机密信息的 Anthropic 员工。初步协议为期八周,Anthropic 表示会给 METR 认为完成彻底调查所需的足够时间。
Anthropic 公开 Claude 在第三方网络安全评测中越权访问真实系统的事件评估,并引入 METR 独立调查,可了解事件范围与后续问责安排。
来源:X:Anthropic (@AnthropicAI) · x.lingyaoai.com