用大模型结构化医护对话,减轻临床文书负担。
Empowering Healthcare Practitioners with Language Models: Structuring Speech Transcripts in Two Real-World Clinical Applications
- 设计智能代理生成真实医护语音数据,解决数据稀缺问题。
- 在两个真实医疗场景中实现高精度结构化报告与医嘱提取。
- 开源首个医护观察与医嘱抽取数据集,推动临床NLP研究。
大型语言模型(如GPT-4o和o1)在多个医学自然语言处理基准上表现优异。然而,由于数据稀缺与敏感性,护士口述转录的结构化表格报告及医生-患者对话中的医嘱提取这两项高影响力任务仍鲜有研究,尽管产业界已有积极尝试。本论文针对这两个实际临床任务,利用私有与开源临床数据集,评估了开源与闭源语言模型的表现,并分析其优劣。此外,提出一种智能代理流程,用于生成真实、无敏感信息的护士口述数据,支持临床观察的结构化提取。为促进后续研究,我们发布了SYNUR与SIMORD——首个面向护士观察提取与医嘱提取的开源数据集。
原文摘要 · Abstract (English)
Large language models (LLMs) such as GPT-4o and o1 have demonstrated strong performance on clinical natural language processing (NLP) tasks across multiple medical benchmarks. Nonetheless, two high-impact NLP tasks - structured tabular reporting from nurse dictations and medical order extraction from doctor-patient consultations - remain underexplored due to data scarcity and sensitivity, despite active industry efforts. Practical solutions to these real-world clinical tasks can significantly reduce the documentation burden on healthcare providers, allowing greater focus on patient care. In this paper, we investigate these two challenging tasks using private and open-source clinical datasets, evaluating the performance of both open- and closed-weight LLMs, and analyzing their respective strengths and limitations. Furthermore, we propose an agentic pipeline for generating realistic, non-sensitive nurse dictations, enabling structured extraction of clinical observations. To support further research in both areas, we release SYNUR and SIMORD, the first open-source datasets for nurse observation extraction and medical order extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。