AI代理需警惕内容误导,事实核查可成防御新手段
Attacks by Content: Automated Fact-checking is an AI Security Issue
- 用虚假或误导性信息直接操纵AI行为,无需注入指令
- 现有防御对内容攻击无效,因仅针对隐藏命令检测
- 将自动事实核查技术用于AI自我防御,提升信息可信度判断
当AI代理从外部文档中检索并推理信息时,攻击者可通过提供带有偏见、误导或虚假内容的资料来操控其行为。以往研究关注间接提示注入攻击,即通过植入恶意指令实现控制。本文认为,攻击者无需注入指令,仅通过提供错误内容即可达成操纵目的,称之为‘内容攻击’。现有防御机制聚焦于识别隐藏指令,对这类内容攻击毫无作用。为保护自身与用户,代理必须具备批判性评估能力,包括交叉验证信息真实性、评估信息来源可信度。我们提出,这一能力与自然语言处理中的自动化事实核查任务高度相似,可将其重构为代理的认知自卫工具。
原文摘要 · Abstract (English)
When AI agents retrieve and reason over external documents, adversaries can manipulate the data they receive to subvert their behaviour. Previous research has studied indirect prompt injection, where the attacker injects malicious instructions. We argue that injection of instructions is not necessary to manipulate agents - attackers could instead supply biased, misleading, or false information. We term this an attack by content. Existing defenses, which focus on detecting hidden commands, are ineffective against attacks by content. To defend themselves and their users, agents must critically evaluate retrieved information, corroborating claims with external evidence and evaluating source trustworthiness. We argue that this is analogous to an existing NLP task, automated fact-checking, which we propose to repurpose as a cognitive self-defense tool for agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。