arXiv:2507.01049cs.IRcs.CL2025-07

用密集段落检索技术,从心超病历中精准找患者群体

Cohort Retrieval using Dense Passage Retrieval

  • 将心超电子病历转为查询-段落对,用稠密检索方法建模
  • 自研模型在临床场景下表现优于现有最优方法
  • 首个在心超领域应用DPR的框架,可推广至其他医学领域

患者队列检索是医学研究和临床实践中的关键任务,有助于从海量电子健康记录(EHR)中识别特定患者群体。本文针对超声心动图领域的队列检索挑战,采用密集段落检索(Dense Passage Retrieval, DPR)这一语义搜索主流方法。我们提出一种系统性方法,将非结构化的超声心动图EHR数据转换为查询-段落数据集,并将问题建模为队列检索任务。此外,我们设计并实现了受真实临床场景启发的评估指标,以在多种检索任务中严格测试模型性能。我们还提出一个定制训练的DPR嵌入模型,在多个任务上表现优于传统及现成的先进方法。据我们所知,这是首个在超声心动图领域应用DPR进行患者队列检索的工作,建立了一个可扩展至其他医学领域的框架。

原文摘要 · Abstract (English)

Patient cohort retrieval is a pivotal task in medical research and clinical practice, enabling the identification of specific patient groups from extensive electronic health records (EHRs). In this work, we address the challenge of cohort retrieval in the echocardiography domain by applying Dense Passage Retrieval (DPR), a prominent methodology in semantic search. We propose a systematic approach to transform an echocardiographic EHR dataset of unstructured nature into a Query-Passage dataset, framing the problem as a Cohort Retrieval task. Additionally, we design and implement evaluation metrics inspired by real-world clinical scenarios to rigorously test the models across diverse retrieval tasks. Furthermore, we present a custom-trained DPR embedding model that demonstrates superior performance compared to traditional and off-the-shelf SOTA methods.To our knowledge, this is the first work to apply DPR for patient cohort retrieval in the echocardiography domain, establishing a framework that can be adapted to other medical domains.

医学检索稠密检索队列分析心超数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。