arXiv:2410.23090cs.IRcs.CL2024-10被引 15

构建多轮对话检索增强生成基准,填补真实场景评估空白

CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation

  • 基于维基百科自动构建多轮对话数据集,覆盖开放域与话题跳跃
  • 首次在真实对话场景下评估检索、生成、引用标注三任务性能
  • 提供统一框架与评测标准,适合研究对话增强生成的团队使用

检索增强生成(RAG)已成为通过外部知识检索提升大语言模型能力的重要范式。尽管受到广泛关注,现有学术研究主要集中在单轮RAG,难以应对真实应用中复杂的多轮对话挑战。为弥补这一差距,我们提出CORAL,一个大规模基准,用于评估RAG系统在真实多轮对话场景中的表现。CORAL基于维基百科自动生成多样化的信息查询对话,涵盖开放域覆盖、知识密集性、自由格式回复及话题转移等关键挑战。该基准支持对话RAG的三项核心任务:段落检索、响应生成与引用标注。我们提出统一框架以标准化不同对话RAG方法,并在CORAL上开展全面评估,揭示了现有方法仍有显著提升空间。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) has become a powerful paradigm for enhancing large language models (LLMs) through external knowledge retrieval. Despite its widespread attention, existing academic research predominantly focuses on single-turn RAG, leaving a significant gap in addressing the complexities of multi-turn conversations found in real-world applications. To bridge this gap, we introduce CORAL, a large-scale benchmark designed to assess RAG systems in realistic multi-turn conversational settings. CORAL includes diverse information-seeking conversations automatically derived from Wikipedia and tackles key challenges such as open-domain coverage, knowledge intensity, free-form responses, and topic shifts. It supports three core tasks of conversational RAG: passage retrieval, response generation, and citation labeling. We propose a unified framework to standardize various conversational RAG methods and conduct a comprehensive evaluation of these methods on CORAL, demonstrating substantial opportunities for improving existing approaches.

对话RAG检索增强评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。