arXiv:2412.06954cs.IR2024-12KDD被引 5

构建医疗场景检索评估数据集CURE,支持多语言测试。

CURE: A Dataset for Clinical Understanding & Retrieval Evaluation

  • 联合医生构建2000个查询的临床检索数据集
  • 覆盖10个医学领域,含英/法/西三语条件
  • 可用于评估医疗场景下检索模型性能

由于密集检索器在训练数据分布外泛化能力差,特定领域测试集对检索系统评估至关重要。目前针对临床诊疗场景的检索测试集仍很匮乏。为此,我们与医疗专业人员合作,创建了CURE——一个用于段落排序的即席检索测试数据集,包含2000个查询,覆盖10个医学领域,并设有单语(英语)及双跨语言(法语/西班牙语→英语)条件。本文详述CURE的构建过程,并提供基线结果以展示其作为评估工具的有效性。CURE采用知识共享署名-非商业性4.0许可,可在Hugging Face上获取,并作为MTEB上的检索任务使用。

原文摘要 · Abstract (English)

Given the dominance of dense retrievers that do not generalize well beyond their training dataset distributions, domain-specific test sets are essential in evaluating retrieval. There are few test datasets for retrieval systems intended for use by healthcare providers in a point-of-care setting. To fill this gap we have collaborated with medical professionals to create CURE, an ad-hoc retrieval test dataset for passage ranking with 2000 queries spanning 10 medical domains with a monolingual (English) and two cross-lingual (French/Spanish -> English) conditions. In this paper, we describe how CURE was constructed and provide baseline results to showcase its effectiveness as an evaluation tool. CURE is published with a Creative Commons Attribution Non Commercial 4.0 license and can be accessed on Hugging Face and as a retrieval task on MTEB.

医疗检索数据集跨语言评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。