测试开源大模型对故意误导输入的识别能力,发现自信程度越低反而越难被识破。
Battling Misinformation: An Empirical Study on Adversarial Factuality in Open-Source Large Language Models
- 设计三档信心水平的误导性提示,测试8个开源大模型的抗骗能力
- 多数模型在对手信心弱时更易识别,但Llama 3.1和Phi 3反常
- 针对冷门知识的攻击更有效,提示越冷门越容易成功欺骗
对抗性真实性指攻击者在输入提示中刻意插入虚假信息,且表达不同程度的信心。本研究系统评估了多个开源大语言模型(LLMs)在面对此类对抗性输入时的表现。考虑三种对抗信心层级:高度自信、中等自信和有限自信。分析涵盖八款模型:LLaMA 3.1 (8B)、Phi 3 (3.8B)、Qwen 2.5 (7B)、Deepseek-v2 (16B)、Gemma2 (9B)、Falcon (7B)、Mistralite (7B) 和 LLaVA (7B)。实证结果表明,LLaMA 3.1 (8B) 在检测对抗输入方面表现出较强能力,而 Falcon (7B) 表现相对较低。值得注意的是,多数模型在对抗者信心降低时检测成功率提升;但这一趋势在 LLaMA 3.1 (8B) 与 Phi 3 (3.8B) 上出现反转,即对抗信心减弱反而导致检测性能下降。进一步分析高/低成功率攻击所对应的查询发现,针对较少引用或冷门信息的攻击更具成效。
原文摘要 · Abstract (English)
Adversarial factuality refers to the deliberate insertion of misinformation into input prompts by an adversary, characterized by varying levels of expressed confidence. In this study, we systematically evaluate the performance of several open-source large language models (LLMs) when exposed to such adversarial inputs. Three tiers of adversarial confidence are considered: strongly confident, moderately confident, and limited confidence. Our analysis encompasses eight LLMs: LLaMA 3.1 (8B), Phi 3 (3.8B), Qwen 2.5 (7B), Deepseek-v2 (16B), Gemma2 (9B), Falcon (7B), Mistrallite (7B), and LLaVA (7B). Empirical results indicate that LLaMA 3.1 (8B) exhibits a robust capability in detecting adversarial inputs, whereas Falcon (7B) shows comparatively lower performance. Notably, for the majority of the models, detection success improves as the adversary's confidence decreases; however, this trend is reversed for LLaMA 3.1 (8B) and Phi 3 (3.8B), where a reduction in adversarial confidence corresponds with diminished detection performance. Further analysis of the queries that elicited the highest and lowest rates of successful attacks reveals that adversarial attacks are more effective when targeting less commonly referenced or obscure information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。