简单选择题里藏恶意指令,大模型竟被轻易误导。
Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- 在PDF中嵌入隐藏指令,诱导大模型答错基础算术题。
- 多个主流大模型在简单题目上误判率超50%。
- 适合关注大模型安全与评估可靠性的研究者。
大型语言模型(LLMs)近期在复杂推理和零样本泛化方面展现出强大能力,为教育、同行评审和数据质量评估中的「大模型作为评判者」应用提供了前所未有的潜力。然而,其在提示注入攻击下的鲁棒性仍令人担忧——此类攻击通过将恶意指令嵌入内容来操控输出。本文探索了一种看似简单却极为有效的攻击场景:在包含基本算术题(如“3 + 2 是多少?”)的PDF文件中,以多选或真假判断形式呈现问题,同时在文件中嵌入隐藏提示。结果表明,即使在这些微不足道的任务中,大模型仍极易受到隐藏提示注入攻击的影响,暴露出「大模型作为评判者」应用中的严重鲁棒性风险。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently demonstrated strong emergent abilities in complex reasoning and zero-shot generalization, showing unprecedented potential for LLM-as-a-judge applications in education, peer review, and data quality evaluation. However, their robustness under prompt injection attacks, where malicious instructions are embedded into the content to manipulate outputs, remains a significant concern. In this work, we explore a frustratingly simple yet effective attack setting to test whether LLMs can be easily misled. Specifically, we evaluate LLMs on basic arithmetic questions (e.g., "What is 3 + 2?") presented as either multiple-choice or true-false judgment problems within PDF files, where hidden prompts are injected into the file. Our results reveal that LLMs are indeed vulnerable to such hidden prompt injection attacks, even in these trivial scenarios, highlighting serious robustness risks for LLM-as-a-judge applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。