arXiv:2507.05713cs.CLcs.AI2025-07Conference of the …被引 7

构建定期更新的RAG评估基准,解决数据泄露问题。

DRAGOn: Designing RAG On Periodically Updated Corpus

  • 基于知识图谱自动生成问答对,支持动态数据更新。
  • 使用俄语新闻源构建数据集,确保模型评测在未见数据上进行。
  • 提供自动评估与公开排行榜,促进社区协作与进步。

本文提出DRAGOn,一种在定期更新语料库上构建RAG评估基准的方法。其包含最新参考数据集、问答生成框架、自动评估流水线及公开排行榜。指定参考数据集实现RAG系统间的统一比较,新生成的数据版本可避免数据泄露,确保所有模型在未见且可比的数据上测试。问答生成管道从文本语料中提取知识图谱,并利用现代大模型能力生成多组问答对。评估采用多样化的LLM-as-Judge指标,实现全面评价。研究以俄语新闻媒体为数据源,验证方法有效性,并上线公开排行榜,推动RAG技术发展与社区参与。

原文摘要 · Abstract (English)

This paper introduces DRAGOn, method to design a RAG benchmark on a regularly updated corpus. It features recent reference datasets, a question generation framework, an automatic evaluation pipeline, and a public leaderboard. Specified reference datasets allow for uniform comparison of RAG systems, while newly generated dataset versions mitigate data leakage and ensure that all models are evaluated on unseen, comparable data. The pipeline for automatic question generation extracts the Knowledge Graph from the text corpus and produces multiple question-answer pairs utilizing modern LLM capabilities. A set of diverse LLM-as-Judge metrics is provided for a comprehensive model evaluation. We used Russian news outlets to form the datasets and demonstrate our methodology. We launch a public leaderboard to track the development of RAG systems and encourage community participation.

RAG评估基准知识图谱自动评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。