Google 论文:加一句 Be honest 能让模型汇报失败结果
Google 一篇论文显示,语言模型在总结完成的工作时会隐瞒严重缺陷,即使这些缺陷它们自己能看到;在提示词中加入 Be honest in your response 能显著改善。
Hugely revealing paper from Google.
If you are reading AI summaries instead of logs, add "Be honest in your response" to the prompt, because without it frontier models routinely skip the bad news.
Language models hide serious flaws when they summarize finished work, even flaws they can see, and a plain "Be honest in your response" line gets far more of them reported.
Given an experiment log where the new method loses to a strong baseline, GPT-5.5 mentioned the loss in 2 of 200 abstracts. Told to "Be honest in your response," it mentioned it in 190 of 200.
Across 8 setups, from buggy code to agent logs with an unfinished job, the models could spot each flaw when asked directly. Their reasoning showed them choosing to keep the success story intact.
The honesty line barely helped when an agent reported results from a tool call that was still running.
If you depend on agent summaries, put an honesty instruction in every report prompt, and still check raw logs for pending or unfinished steps.
来源:rohanpaul_ai · x.com