让医生用自然语言快速查询病历数据,准确率更高。
Generating Querying Code from Text for Multi-Modal Electronic Health Record
- 用医学知识模块和模板匹配技术,把中文问题转为结构化查询语句。
- 在自建数据集TQGen上,查询准确率达78.3%,显著优于基线模型。
- 适合医疗信息化、临床决策支持系统研发人员参考。
电子健康记录(EHR)包含大量结构化与非结构化数据,如表格信息和自由文本临床笔记。提取相关信息常需复杂的数据库操作,增加临床工作负担。然而,复杂的表关系和专业术语限制了查询准确性。本文构建了一个公开数据集TQGen,整合了表格与临床文本,用于自然语言到查询生成任务。为应对医学术语复杂性和问题多样性,提出TQGen-EHRQuery框架,包含医学知识模块与问题模板匹配模块。针对医学文本处理,引入工具集概念,将文本处理模块封装为可调用工具,提升处理效率与灵活性。通过大量实验评估数据集与流程有效性,验证其在提升EHR信息查询方面的潜力。
原文摘要 · Abstract (English)
Electronic health records (EHR) contain extensive structured and unstructured data, including tabular information and free-text clinical notes. Querying relevant patient information often requires complex database operations, increasing the workload for clinicians. However, complex table relationships and professional terminology in EHRs limit the query accuracy. In this work, we construct a publicly available dataset, TQGen, that integrates both \textbf{T}ables and clinical \textbf{T}ext for natural language-to-query \textbf{Gen}eration. To address the challenges posed by complex medical terminology and diverse types of questions in EHRs, we propose TQGen-EHRQuery, a framework comprising a medical knowledge module and a questions template matching module. For processing medical text, we introduced the concept of a toolset, which encapsulates the text processing module as a callable tool, thereby improving processing efficiency and flexibility. We conducted extensive experiments to assess the effectiveness of our dataset and workflow, demonstrating their potential to enhance information querying in EHR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。