arXiv:2503.09899cs.IR2025-03被引 7

用大模型填补对话搜索测试集的空白文档,提升新系统评估公平性。

Improving the Reusability of Conversational Search Test Collections

  • 用Llama 3.1少样本微调填充未标注文档,效果更接近人工判断。
  • 对话搜索测试集在深层交互中可复用性下降,空白问题更严重。
  • 该方法让新系统与旧系统对比更公平,适合评测新对话系统的研究者。

不完整的相关性判断限制了测试集的可复用性。当新系统与曾贡献过结果的旧系统比较时,常因测试集中存在未标注文档(称为‘空洞’)而处于不利地位。对话搜索(CS)的特性使得这些空洞更大、更影响评估。本文使用大语言模型(LLMs)填补空洞,基于已有判断进行扩展。实验以TREC iKAT 23和TREC CAsT 22数据集为基础,信息需求动态性强,响应多样,空洞较大。结果显示,深度交互轮次下测试集可复用性下降。微调Llama 3.1模型与人工评估高度一致,而使用ChatGPT少样本提示的结果一致性低。若用ChatGPT生成评估池,新系统排名变化显著;但采用其少样本提示重生成评估池,与人工池的排名相关性高。使用少样本微调的Llama 3.1填充空洞,可实现新系统与旧系统的公平比较,有效提升测试集可复用性。

原文摘要 · Abstract (English)

Incomplete relevance judgments limit the reusability of test collections. When new systems are compared to previous systems that contributed to the pool, they often face a disadvantage. This is due to pockets of unjudged documents (called holes) in the test collection that the new systems return. The very nature of Conversational Search (CS) means that these holes are potentially larger and more problematic when evaluating systems. In this paper, we aim to extend CS test collections by employing Large Language Models (LLMs) to fill holes by leveraging existing judgments. We explore this problem using TREC iKAT 23 and TREC CAsT 22 collections, where information needs are highly dynamic and the responses are much more varied, leaving bigger holes to fill. Our experiments reveal that CS collections show a trend towards less reusability in deeper turns. Also, fine-tuning the Llama 3.1 model leads to high agreement with human assessors, while few-shot prompting the ChatGPT results in low agreement with humans. Consequently, filling the holes of a new system using ChatGPT leads to a higher change in the location of the new system. While regenerating the assessment pool with few-shot prompting the ChatGPT model and using it for evaluation achieves a high rank correlation with human-assessed pools. We show that filling the holes using few-shot training the Llama 3.1 model enables a fairer comparison between the new system and the systems contributed to the pool. Our hole-filling model based on few-shot training of the Llama 3.1 model can improve the reusability of test collections.

对话搜索测试集大模型可复用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。