rohanpaul_ai· @rohanpaul_ai · X·· 2 天前AI 评分
Anthropic:模型自述推理不可信
Anthropic 明确表示,模型对自己推理过程的解释不能作为其行为原因的证据,这也正是他们无法准确判断这些故障严重程度的原因。
Anthropic states plainly that the model’s own explanation of its reasoning can’t be trusted as evidence of why it acted, which is exactly why they can’t cleanly judge how severe each of these failures was.
来源:rohanpaul_ai · x.com