攻击者可骗大模型记错事实,且成功率高达85%
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
- 用理论驱动的中间人框架,通过指令干扰让大模型答错事实
- 简单指令攻击成功率达85.3%,错误回答时不确定性极高
- 基于响应不确定性的分类器可94.8%识别攻击,适合安全防护
大模型已成为信息检索的重要组成部分,但其作为问答聊天机器人的角色面临严重挑战,因其易受对抗性中间人(MitM)攻击。本文提出首个针对大模型事实记忆的系统性攻击评估方法——Xmera,一种基于理论的新型MitM框架。在三个闭卷、基于事实的问答设置中,通过扰动输入,我们削弱了大模型回答的正确性,并评估其生成过程的不确定性。令人惊讶的是,简单的指令型攻击成功率最高,可达约85.3%,且错误回答时具有高不确定性。为提供简易防御机制,我们利用响应不确定性训练随机森林分类器,以区分被攻击与未被攻击查询,平均AUC达约94.8%。我们认为,向用户警示从黑箱可能被污染的大模型获取答案的风险,是保障网络空间安全的第一步。
原文摘要 · Abstract (English)
LLMs are now an integral part of information retrieval. As such, their role as question answering chatbots raises significant concerns due to their shown vulnerability to adversarial man-in-the-middle (MitM) attacks. Here, we propose the first principled attack evaluation on LLM factual memory under prompt injection via Xmera, our novel, theory-grounded MitM framework. By perturbing the input given to "victim" LLMs in three closed-book and fact-based QA settings, we undermine the correctness of the responses and assess the uncertainty of their generation process. Surprisingly, trivial instruction-based attacks report the highest success rate (up to ~85.3%) while simultaneously having a high uncertainty for incorrectly answered questions. To provide a simple defense mechanism against Xmera, we train Random Forest classifiers on the response uncertainty levels to distinguish between attacked and unattacked queries (average AUC of up to ~94.8%). We believe that signaling users to be cautious about the answers they receive from black-box and potentially corrupt LLMs is a first checkpoint toward user cyberspace safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。