arXiv:2606.26879cs.AI2026-06

用大模型生成真实感临床记录,助力医疗AI研发

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

论文配图:A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models
图 1 · 摘自论文原文
  • 分三步生成:患者信息→病程模拟→病历文本生成
  • 70名虚拟患者,每人20-50份病历,覆盖完整住院流程
  • 支持多种验证级别,适合不同场景的医疗AI测试

合成数据正被广泛用于受限数据领域的AI系统开发与评估。在医疗领域,临床文档因敏感性面临特殊挑战。本文提出一个合成临床病历生成流水线及数据集,旨在支持临床AI工具开发,同时规避真实患者数据的隐私风险。该数据集通过模块化流程生成:结构化患者信息、半结构化病程模拟,以及基于大语言模型的非结构化病历生成。流程注重纵向病历内部一致性,同时保留书写风格、格式和临床细节的多样性。引入基于LLM的验证与增强机制,提升生成内容的真实性、合理性和多样性。我们发布包含70名合成患者的数据集,每名患者关联20至50份病历,涵盖完整医院就诊过程。数据提供多级验证版本,用户可根据需求平衡真实性与可扩展性。该数据集可用于临床AI系统的开发、测试与评估,包括摘要生成、编码模型和决策支持系统,无需依赖真实患者数据。

原文摘要 · Abstract (English)

Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.

合成数据临床病历大模型应用医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。