arXiv:2603.17176cs.CRcs.AI2026-03

无需标注数据,自动检测生成系统中的恶意文档。

Towards Unsupervised Adversarial Document Detection in Retrieval Augmented Generation Systems

  • 利用生成激活、输出嵌入和熵不确定性作为异常指标
  • 在不依赖目标提示的情况下实现零日攻击检测
  • 简单摘要生成可能比复杂模型更有效

检索增强生成系统已广泛应用于搜索引擎、邮件系统及服务聊天机器人中,基于大语言模型进行上下文检索与答案生成。随着其普及,安全漏洞日益突出,攻击者通过篡改上下文文档实施持久性攻击,影响所有用户。因此,早期识别受污染的对抗性上下文至关重要。现有监督方法依赖大量标注数据,本文提出一种无监督方法,可检测零日攻击。我们初步研究发现,生成器激活、输出嵌入和基于熵的不确定性度量是有效的互补指标。通过基础统计异常检测方法评估其性能,结果表明:无需目标提示即可成功检测,且简单的上下文摘要生成在识别篡改内容方面表现更优。

原文摘要 · Abstract (English)

Retrieval augmented generation systems have become an integral part of everyday life. Whether in internet search engines, email systems, or service chatbots, these systems are based on context retrieval and answer generation with large language models. With their spread, also the security vulnerabilities increase. Attackers become increasingly focused on these systems and various hacking approaches are developed. Manipulating the context documents is a way to persist attacks and make them affect all users. Therefore, detecting compromised, adversarial context documents early is crucial for security. While supervised approaches require a large amount of labeled adversarial contexts, we propose an unsupervised approach, being able to detect also zero day attacks. We conduct a preliminary study to show appropriate indicators for adversarial contexts. For that purpose generator activations, output embeddings, and an entropy-based uncertainty measure turn out as suitable, complementary quantities. With an elementary statistical outlier detection, we propose and compare their detection abilities. Furthermore, we show that the target prompt, which the attacker wants to manipulate, is not required for a successful detection. Moreover, our results indicate that a simple context summary generation might even be superior in finding manipulated contexts.

文档检测无监督学习安全防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。