arXiv:2608.04366cs.CRcs.AI2026-08被引 4

对抗大模型知识污染,通过可信验证确保检索增强生成安全

Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework

论文配图:Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework
图 1 · 摘自论文原文
  • 多源知识验证+动态图神经网络评分,识别伪造文档来源
  • 在非独立同分布数据下仍能抵御攻击,保持知识完整性
  • 适合部署在敏感场景的协作式AI系统,如金融、医疗

尽管检索增强生成系统部分缓解了大语言模型的幻觉问题,但也引入了知识污染攻击的新漏洞。攻击者通过污染RAG系统提供的文档来操纵LLM输出。为应对这一威胁,我们提出SecureCollaRAG,一种基于多源知识验证机制的拜占庭容错协作RAG框架。该方法利用动态图神经网络进行可信度评分,实现对文档来源的安全验证,有效防范隐蔽的知识污染攻击,同时保留关键领域知识完整性。通过大量实验与形式化分析,我们证明SecureCollaRAG在非独立同分布(non-IID)数据分布下仍具备强鲁棒性。

原文摘要 · Abstract (English)

While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisoning documents provided by RAG system to manipulate LLM outputs. To counter this threat, we propose SecureCollaRAG, a Byzantine-tolerant collaborative RAG framework leveraging Multi-source Knowledge Validation Mechanism. Our approach enables agent system to securely verify document provenance through dynamic GNN-based credibility scoring, effectively preventing stealthy knowledge corruption attacks while preserving essential domain knowledge integrity. Through extensive evaluations and formal analysis, we demonstrate that SecureCollaRAG maintains robustness against attackers under non-IID data distributions.

RAG安全知识污染图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。