用合成数据训练临床沟通模型,解决真实数据稀缺难题
Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies
- 将病历等结构化数据转为自然语言通信用于训练
- 合成数据训练的模型在多个场景表现媲美真实数据模型
- 适合无标注数据的临床场景如急诊、交接班等研究者
大量临床价值存在于非结构化沟通中:患者描述症状、医生推理下达指令、救护车交接、护士换班传递信息等。这些语言依赖角色、意图、因果、不确定性、省略和信道噪声,无法通过表格数据建模。医疗自然语言处理需理解实际传达的信息而非编码内容。但高质量标注语料稀缺,因真实交流涉及隐私、碎片化且标注成本高。大语言模型可通过转化病历、诊断标签、症状列表或护理计划等结构化数据,生成可用于下游任务的书面与语音通信。本文按数据源、沟通形式、参与者、生成方法和下游任务进行系统综述,并提供13个新案例研究。涵盖急救前报告、现场无线电伤员记录、护士交接班、患者门户分诊及低资源出院沟通等无真实标注数据的场景,验证了合成通信可驱动系统构建。结果表明微调编码器模型优于零样本基线,刻意降质的通信有助于提升鲁棒性。主要局限在于多数评估基于合成数据内部测试,缺乏训练用合成、测试用真实数据的验证。结论是合成临床沟通正成为可行研究资源;要使其成为可复用的临床基础设施,还需真实数据迁移、安全性和外部验证。
原文摘要 · Abstract (English)
Much clinical value is conveyed not through structured records but through communication: exchanges in which patients describe symptoms, clinicians reason and give instructions, ambulances hand over to emergency departments, and nurses pass on a shift. Such language differs from tabular data because meaning depends on speaker role, intent, causality, uncertainty, omission, and channel noise. Healthcare natural language processing must therefore interpret information as conveyed rather than coded. This requires well-annotated corpora, which are scarce because authentic exchanges are private, fragmented, and costly to annotate. Large language models offer a way forward by transforming clinical sources, such as records, diagnostic labels, symptom lists, or care plans, into written and transcribed communication for downstream models. We present a structured narrative survey organized by source representation, communication form and participants, generation method, and downstream task, complemented by thirteen novel case studies. These build clinical NLP systems for communication channels and languages without labeled real-world data, including EMS pre-arrival reports, field-radio casualty documentation, nurse handoffs, patient-portal triage, and low-resource discharge communication. They show that synthetic communication can bootstrap such systems. Findings include the competitiveness of fine-tuned encoder models over evaluated zero-shot baselines and the value of deliberately degraded communication for robustness. The main limitation is that most studies evaluate on held-out synthetic communication, while train-on-synthetic, test-on-authentic evidence remains limited. We conclude that syn-thetic clinical communication is becoming a practical research resource; establishing it as reusable clinical infrastructure will require authentic-data transfer, safety and external validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。