为RAG系统构建安全风险评估与缓解框架,保障敏感数据不泄露。
Securing RAG: A Risk Assessment and Mitigation Framework
- 梳理RAG全流程漏洞,覆盖数据预处理到大模型集成
- 提出融合行业标准的安全框架,提升系统可信度
- 适合部署RAG的工程师和安全团队参考
检索增强生成(RAG)已成为面向用户的NLP应用的行业标准,能够在不重新训练或微调大语言模型(LLMs)的情况下整合数据,从而提升响应质量与准确性。然而,这一能力也引入了新的安全与隐私挑战,尤其在敏感数据被集成时更为显著。随着RAG的快速普及,保障数据与服务安全已成为关键优先事项。本文首先回顾RAG管道中的漏洞,从数据预处理、存储管理到与LLMs的集成,全面分析攻击面。随后,将识别出的风险与相应的缓解措施进行结构化匹配。第二步,提出一个结合RAG特定安全考量、现有通用安全指南、行业标准及最佳实践的综合框架,旨在指导可信赖、合规、安全的RAG系统实现。
原文摘要 · Abstract (English)
Retrieval Augmented Generation (RAG) has emerged as the de facto industry standard for user-facing NLP applications, offering the ability to integrate data without re-training or fine-tuning Large Language Models (LLMs). This capability enhances the quality and accuracy of responses but also introduces novel security and privacy challenges, particularly when sensitive data is integrated. With the rapid adoption of RAG, securing data and services has become a critical priority. This paper first reviews the vulnerabilities of RAG pipelines, and outlines the attack surface from data pre-processing and data storage management to integration with LLMs. The identified risks are then paired with corresponding mitigations in a structured overview. In a second step, the paper develops a framework that combines RAG-specific security considerations, with existing general security guidelines, industry standards, and best practices. The proposed framework aims to guide the implementation of robust, compliant, secure, and trustworthy RAG systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。