研究大模型自信表达在对抗攻击下的脆弱性,发现小改动就能骗过模型自信判断。
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
- 设计两类攻击:扰动和越狱,破坏模型口头自信评分
- 攻击后模型自信估计严重失真,常导致答案频繁切换
- 现有防御手段基本无效甚至适得其反,适合关注AI可信性的研究者
大型语言模型(LLMs)生成的稳健口头自信对于确保人机交互等应用中的透明度、信任与安全至关重要。本文首次系统研究了在对抗攻击下口头自信的鲁棒性。我们提出了针对口头自信分数的攻击框架,涵盖扰动与越狱两种方法,证明这些攻击能显著削弱自信评估并引发频繁的答案变更。通过考察多种提示策略、模型规模与应用场景,发现当前口头自信普遍脆弱,且常用防御手段大多无效或适得其反。研究强调需设计更稳健的自信表达机制,因为即使细微的语义保持修改也可能导致响应自信误导。
原文摘要 · Abstract (English)
Robust verbal confidence generated by large language models (LLMs) is crucial for the deployment of LLMs to help ensure transparency, trust, and safety in many applications, including those involving human-AI interactions. In this paper, we present the first comprehensive study on the robustness of verbal confidence under adversarial attacks. We introduce attack frameworks targeting verbal confidence scores through both perturbation and jailbreak-based methods, and demonstrate that these attacks can significantly impair verbal confidence estimates and lead to frequent answer changes. We examine a variety of prompting strategies, model sizes, and application domains, revealing that current verbal confidence is vulnerable and that commonly used defence techniques are largely ineffective or counterproductive. Our findings underscore the need to design robust mechanisms for confidence expression in LLMs, as even subtle semantic-preserving modifications can lead to misleading confidence in responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。