arXiv:2504.15689cs.IR2025-04中稿 · SIGIR'25被引 10

用众包评估RAG效果,发现人类判断更可靠且便宜。

The Viability of Crowdsourcing for RAG Evaluation

  • 构建众包RAG语料库,含903条人工与903条AI生成回复。
  • 47,320次人工两两对比判断显示人类评价更可信。
  • 适合研究RAG评估方法或想低成本获取高质量判断的研究者。

为评估人类在检索增强生成(RAG)场景中撰写和评判回答的能力,我们通过两项互补研究考察了众包在RAG评估中的有效性:回答生成与回答效用判断。我们构建了Crowd RAG Corpus 2025(CrowdRAG-25),包含TREC RAG'24赛道301个主题的903条人工撰写与903条大模型生成的回答,涵盖“要点列表”“论文”“新闻”三种文体。针对其中65个主题,还包含47,320次人工两两判断与10,556次模型两两判断,覆盖七项效用维度(如覆盖度、连贯性)。分析表明,相比基于模型的点判或人/模型混合评分,人工两两判断在可靠性与成本上更具优势,且优于自动化对比人工参考答案的结果。所有数据与工具均开源可用。

原文摘要 · Abstract (English)

How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RAG through two complementary studies: response writing and response utility judgment. We present the Crowd RAG Corpus 2025 (CrowdRAG-25), which consists of 903 human-written and 903 LLM-generated responses for the 301 topics of the TREC RAG'24 track, across the three discourse styles 'bulleted list', 'essay', and 'news'. For a selection of 65 topics, the corpus further contains 47,320 pairwise human judgments and 10,556 pairwise LLM judgments across seven utility dimensions (e.g., coverage and coherence). Our analyses give insights into human writing behavior for RAG and the viability of crowdsourcing for RAG evaluation. Human pairwise judgments provide reliable and cost-effective results compared to LLM-based pairwise or human/LLM-based pointwise judgments, as well as automated comparisons with human-written reference responses. All our data and tools are freely available.

RAG评估众包人工判断数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。