构建首个多场景电子病历检索基准,助力临床信息精准获取。
CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment
- 基于1000份出院记录与大模型标注,构建双场景检索数据集
- 揭示通用稠密检索器在医疗场景中表现优于专用模型
- 创新性划分语义匹配类型,量化语义鸿沟问题
电子病历(EHR)检索在临床任务中至关重要,但其发展受限于缺乏公开基准。本文提出新基准 ClinIQ,涵盖单患者与多患者两种检索场景:前者聚焦病历内部相关片段定位,后者涉及跨患者病历检索。基准基于 MIMIC-III 的 1,000 份出院记录,包含 ICD 编码与处方标签,并利用大模型生成 1,246 个唯一查询及 77,206 条相关性判断。引入新颖的语义匹配评估机制,将匹配类型细分为字符串匹配和四类语义匹配。在该基准上,对从精确匹配到主流稠密检索器的多种方法进行综合评估。实验发现,BM25 设定强基线,性能媲美稠密检索器;通用领域稠密检索器意外优于专为医疗设计的模型。对各类匹配类型的深入分析揭示了不同方法的优劣,为针对性改进提供方向。我们相信该基准将推动 EHR 检索系统的发展。
原文摘要 · Abstract (English)
Electronic Health Record (EHR) retrieval plays a pivotal role in various clinical tasks, but its development has been severely impeded by the lack of publicly available benchmarks. In this paper, we introduce a novel public EHR retrieval benchmark, CliniQ, to address this gap. We consider two retrieval settings: Single-Patient Retrieval and Multi-Patient Retrieval, reflecting various real-world scenarios. Single-Patient Retrieval focuses on finding relevant parts within a patient note, while Multi-Patient Retrieval involves retrieving EHRs from multiple patients. We build our benchmark upon 1,000 discharge summary notes along with the ICD codes and prescription labels from MIMIC-III, and collect 1,246 unique queries with 77,206 relevance judgments by further leveraging powerful LLMs as annotators. Additionally, we include a novel assessment of the semantic gap issue in EHR retrieval by categorizing matching types into string match and four types of semantic matches. On our proposed benchmark, we conduct a comprehensive evaluation of various retrieval methods, ranging from conventional exact match to popular dense retrievers. Our experiments find that BM25 sets a strong baseline and performs competitively to the dense retrievers, and general domain dense retrievers surprisingly outperform those designed for the medical domain. In-depth analyses on various matching types reveal the strengths and drawbacks of different methods, enlightening the potential for targeted improvement. We believe that our benchmark will stimulate the research communities to advance EHR retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。