arXiv:2501.12789cs.CLcs.IR2025-01被引 20

用DataMorgana生成多样化问答数据,更好评估RAG系统在真实场景表现

Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana

  • 通过两阶段轻量流程自定义用户和问题类别分布,生成可定制的合成数据
  • 在领域和通用知识语料上均实现词汇、语法、语义层面的显著多样性提升
  • 适合参与SIGIR'25 LiveRAG挑战的研究者,用于构建更贴近真实流量的评测基准

评估检索增强生成(RAG)系统,尤其在特定领域场景中,需要满足应用需求的评测基准。由于真实数据难以获取,通常采用大模型生成合成数据。现有方法为通用型:给定文档生成一个问题以构成问答对。然而,尽管生成的问题个体质量尚可,但缺乏足够多样性,无法覆盖真实用户与RAG系统交互的多种方式。本文提出DataMorgana,一种可高度定制且多样化的合成问答基准生成工具,专为RAG应用设计。该工具支持对用户类型和问题类别及其分布进行精细配置,采用轻量级两阶段流程,确保高效迭代并反映预期流量特征。通过一系列实验,定量与定性验证DataMorgana在领域特定及通用知识语料上均优于现有工具,在词汇、句法、语义层面生成了更丰富的问答集。DataMorgana将作为首批测试版本提供给研究社区,用于即将于2025年2月初宣布的SIGIR'25 LiveRAG挑战赛。

原文摘要 · Abstract (English)

Evaluating Retrieval-Augmented Generation (RAG) systems, especially in domain-specific contexts, requires benchmarks that address the distinctive requirements of the applicative scenario. Since real data can be hard to obtain, a common strategy is to use LLM-based methods to generate synthetic data. Existing solutions are general purpose: given a document, they generate a question to build a Q&A pair. However, although the generated questions can be individually good, they are typically not diverse enough to reasonably cover the different ways real end-users can interact with the RAG system. We introduce here DataMorgana, a tool for generating highly customizable and diverse synthetic Q&A benchmarks tailored to RAG applications. DataMorgana enables detailed configurations of user and question categories and provides control over their distribution within the benchmark. It uses a lightweight two-stage process, ensuring efficiency and fast iterations, while generating benchmarks that reflect the expected traffic. We conduct a thorough line of experiments, showing quantitatively and qualitatively that DataMorgana surpasses existing tools and approaches in producing lexically, syntactically, and semantically diverse question sets across domain-specific and general-knowledge corpora. DataMorgana will be made available to selected teams in the research community, as first beta testers, in the context of the upcoming SIGIR'2025 LiveRAG challenge to be announced in early February 2025.

RAG评估合成数据问答生成数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。