arXiv:2502.21107cs.CL2025-02被引 3

用大模型自动把医学条件转成SQL,精准找患者队列

Generating patient cohorts from electronic health records using two-step retrieval-augmented text-to-SQL generation

  • 两阶段检索增强生成,结合医学知识库解析复杂条件
  • 在EHR数据上达0.75的F1分数,准确捕捉时间与逻辑关系
  • 适合临床研究者快速构建研究队列,节省人工编码时间

临床队列定义对患者招募和观察性研究至关重要,但将纳入/排除标准转化为SQL查询仍具挑战且依赖人工。本文提出一种自动化系统,利用大语言模型结合条件解析、两级检索增强生成、医学概念标准化及SQL生成,从电子健康记录中检索患者队列并构建患者流程。该系统在EHR数据上的队列识别任务中取得0.75的F1分数,有效捕获复杂的时序与逻辑关系。结果表明,该方法在流行病学研究中具备自动化队列生成的可行性。

原文摘要 · Abstract (English)

Clinical cohort definition is crucial for patient recruitment and observational studies, yet translating inclusion/exclusion criteria into SQL queries remains challenging and manual. We present an automated system utilizing large language models that combines criteria parsing, two-level retrieval augmented generation with specialized knowledge bases, medical concept standardization, and SQL generation to retrieve patient cohorts with patient funnels. The system achieves 0.75 F1-score in cohort identification on EHR data, effectively capturing complex temporal and logical relationships. These results demonstrate the feasibility of automated cohort generation for epidemiological research.

医疗AI自然语言转SQL电子病历

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。