大模型难以识别自身被恶意预填充攻击,自述可信度低。
Can LLMs Reliably Self-Report Adversarial Prefills, and How?

- 通过分析安全任务中的自我反思信号,发现模型对攻击响应缺乏识别能力。
- 平均25.3%的模型误认为恶意输入是自己主动生成的意图。
- 攻击者可利用模型自述不靠谱的特点,反向提升攻击成功率。
先前研究显示大语言模型在常规任务中具备不同程度的内省能力。本文将其拓展至安全场景,考察模型能否可靠识别自身输出是否由对抗性预填充攻击诱导。在涵盖3B至70B参数的十款开源指令微调模型及四个安全基准测试中,无一模型能可靠识别其输出已被攻破,模型在预填充响应上声称存在意图的平均比例达25.3%。内省信号主要源于对安全性的推理与拒绝行为。将模型权重正交化于拒绝方向后,预填充与自然输出的声称率差距趋近于零,但该方向并非唯一中介。将问题重构为‘内部意图’与‘外部篡改’时,同一模型产生质异反应。训练模型模仿正确内省答案或优化内省目标虽可提升识别准确率,但此类训练无法迁移到篡改探测任务,反而在多数模型上反向提升了对抗预填充攻击的成功率,仅实现部分缓解。这些发现揭示了安全情境下内省信号的作用机制,并警示了大模型自述可靠性风险。
原文摘要 · Abstract (English)
Prior work shows that large language models (LLMs) exhibit varying degrees of introspective capability on benign tasks. We extend the question to safety contexts and examine how reliably a model can recognize that its own prior response was elicited by an adversarial prefill attack. Across ten open-weight instruction-tuned LLMs from 3B to 70B parameters and four safety benchmarks, no model reliably recognizes its own compromised outputs, with models claiming intent on prefilled responses at an average rate of 25.3%. Introspective signal stems primarily from reasoning about safety and refusal. Orthogonalizing models' weights against the refusal direction collapses the gap between claim rates on prefilled and natural outputs to near zero, though the direction is not its unique mediator. Framing the question as internal intention versus external tampering elicits qualitatively different responses on the same models. Training models to mimic correct introspective answers or optimize an introspective objective can improve the accuracy of introspection, but such training does not transfer to the tampering probe and counterintuitively raises attack success rate under adversarial prefill on most models, amounting to a partial mitigation. These findings outline mechanisms underpinning the observed introspective signals in safety contexts and highlight risks in the reliability of LLM self-reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。