arXiv:2508.18929cs.CLcs.AI2025-08被引 1

生成兼顾多样性和隐私保护的合成数据集,用于更真实可靠的RAG系统评估。

Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework

  • 用多智能体框架分步生成覆盖广、语义丰富的问答对。
  • 在多个领域实现敏感信息有效遮蔽,保障数据隐私。
  • 适合关注AI安全与合规的评估人员使用。

检索增强生成(RAG)系统通过引入外部知识提升大语言模型输出质量,使其更具上下文感知能力。然而,这些系统的有效性与可信度高度依赖于评估方式,尤其是否能反映真实场景中的隐私保护需求。尽管现有研究多聚焦于性能指标设计,但对评估数据集的质量和构建方法关注不足,而数据集正是实现可靠评估的关键。本文提出一种新型多智能体框架,用于生成面向RAG评估的合成问答数据集,强调语义多样性和隐私保护。该框架包含:(1) 多样性智能体,利用聚类技术最大化主题覆盖率与语义变异性;(2) 隐私智能体,跨多个领域检测并遮蔽敏感信息;(3) 问答整理智能体,生成符合隐私要求且具备多样性的高质量问答对作为评估基准。大量实验表明,所生成数据集在多样性上优于基线方法,并在特定领域数据集上实现稳健的隐私遮蔽效果。本工作为更安全、全面的RAG系统评估提供了一条切实可行且符合伦理的道路,也为未来契合不断演进的AI监管与合规标准的评估体系奠定基础。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) systems improve large language model outputs by incorporating external knowledge, enabling more informed and context-aware responses. However, the effectiveness and trustworthiness of these systems critically depends on how they are evaluated, particularly on whether the evaluation process captures real-world constraints like protecting sensitive information. While current evaluation efforts for RAG systems have primarily focused on the development of performance metrics, far less attention has been given to the design and quality of the underlying evaluation datasets, despite their pivotal role in enabling meaningful, reliable assessments. In this work, we introduce a novel multi-agent framework for generating synthetic QA datasets for RAG evaluation that prioritize semantic diversity and privacy preservation. Our approach involves: (1) a Diversity agent leveraging clustering techniques to maximize topical coverage and semantic variability, (2) a Privacy Agent that detects and mask sensitive information across multiple domains and (3) a QA curation agent that synthesizes private and diverse QA pairs suitable as ground truth for RAG evaluation. Extensive experiments demonstrate that our evaluation sets outperform baseline methods in diversity and achieve robust privacy masking on domain-specific datasets. This work offers a practical and ethically aligned pathway toward safer, more comprehensive RAG system evaluation, laying the foundation for future enhancements aligned with evolving AI regulations and compliance standards.

RAG评估合成数据隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。