arXiv:2503.09902cs.IR2025-03被引 23

构建对话搜索评估新数据集,用精华信息点提升生成质量评测

Conversational Gold: Evaluating Personalized Conversational Search System using Gold Nuggets

  • 以精华信息点为核心,建立可自动评估的对话生成框架
  • 包含2279个精华信息点和62条人工标注答案,支持长文本生成评测
  • 适合研究对话搜索、幻觉检测与个性化问答的学者使用

个性化对话搜索系统因大语言模型的发展而兴起,能够为复杂信息需求检索并生成回答。然而,检索增强生成(RAG)系统响应的自动评估仍缺乏研究。本文基于TREC iKAT 2023,扩展至2024年数据集,包含17组对话、20,575条相关段落评估、2,279个提取的黄金信息点,以及由NIST评估员手动撰写的62条黄金答案。该数据集引入五大改进:(1)黄金信息点——从相关段落中提取的简洁核心信息;(2)人工撰写答案作为评估标准;(3)不可回答问题用于评估模型幻觉;(4)更丰富的用户角色设定;(5)从个人文本知识库排序转向分类与选择。基于此资源,我们提出一种长文本生成评估框架,包含信息点提取与匹配,关联检索过程。该资源公开可用,助力个性化对话搜索与长文本生成研究。

原文摘要 · Abstract (English)

The rise of personalized conversational search systems has been driven by advancements in Large Language Models (LLMs), enabling these systems to retrieve and generate answers for complex information needs. However, the automatic evaluation of responses generated by Retrieval Augmented Generation (RAG) systems remains an understudied challenge. In this paper, we introduce a new resource for assessing the retrieval effectiveness and relevance of response generated by RAG systems, using a nugget-based evaluation framework. Built upon the foundation of TREC iKAT 2023, our dataset extends to the TREC iKAT 2024 collection, which includes 17 conversations and 20,575 relevance passage assessments, together with 2,279 extracted gold nuggets, and 62 manually written gold answers from NIST assessors. While maintaining the core structure of its predecessor, this new collection enables a deeper exploration of generation tasks in conversational settings. Key improvements in iKAT 2024 include: (1) ``gold nuggets'' -- concise, essential pieces of information extracted from relevant passages of the collection -- which serve as a foundation for automatic response evaluation; (2) manually written answers to provide a gold standard for response evaluation; (3) unanswerable questions to evaluate model hallucination; (4) expanded user personas, providing richer contextual grounding; and (5) a transition from Personal Text Knowledge Base (PTKB) ranking to PTKB classification and selection. Built on this resource, we provide a framework for long-form answer generation evaluation, involving nuggets extraction and nuggets matching, linked to retrieval. This establishes a solid resource for advancing research in personalized conversational search and long-form answer generation. Our resources are publicly available at https://github.com/irlabamsterdam/CONE-RAG.

对话搜索RAG评估黄金信息点长文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。