arXiv:2504.07584cs.IR2025-04被引 2

用提取和生成资源重振旧检索测试集,提升通用性与复用价值。

REANIMATOR: Reanimate Retrieval Test Collections with Extracted and Synthetic Resources

  • 从PDF中解析文本与表格,用大模型生成合成相关性标签。
  • 在TREC-COVID上验证,提升检索增强生成效果,表格显著改善结果。
  • 支持人机协作校验,适合想低成本复用旧数据的研究者。

检索测试集对评估信息检索系统至关重要,但通常缺乏跨任务泛化能力。为克服此局限,我们提出REANIMATOR——一个可复用现有测试集的通用框架,通过引入提取的和合成的资源进行增强。该框架从PDF文件中解析全文、可读取的表格及相关上下文信息,并利用先进的大型语言模型生成合成的相关性标签。可选的人机协同步骤有助于验证所提取与生成的资源。我们在更新后的TREC-COVID测试集上展示了其潜力,构建了检索增强生成系统,并评估了表格对检索增强生成的影响。REANIMATOR实现了测试集在新应用中的再利用,降低开发成本,拓展了遗留资源的使用范围。

原文摘要 · Abstract (English)

Retrieval test collections are essential for evaluating information retrieval systems, yet they often lack generalizability across tasks. To overcome this limitation, we introduce REANIMATOR, a versatile framework designed to enable the repurposing of existing test collections by enriching them with extracted and synthetic resources. REANIMATOR enhances test collections from PDF files by parsing full texts and machine-readable tables, as well as related contextual information. It then employs state-of-the-art large language models to produce synthetic relevance labels. Including an optional human-in-the-loop step can help validate the resources that have been extracted and generated. We demonstrate its potential with a revitalized version of the TREC-COVID test collection, showcasing the development of a retrieval-augmented generation system and evaluating the impact of tables on retrieval-augmented generation. REANIMATOR enables the reuse of test collections for new applications, lowering costs and broadening the utility of legacy resources.

信息检索测试集大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。