arXiv:2509.10538cs.LGcs.AI2025-09

用双重对齐生成真实可信的医疗合成数据,解决隐私与稀有病数据难题。

DualAlign: Generating Clinically Grounded Synthetic Data

  • 通过人口统计与症状轨迹双重约束生成临床合理文本。
  • 在阿尔茨海默病数据上提升模型性能,接近真实标注数据效果。
  • 适合低资源临床文本分析,兼顾隐私保护与数据可用性。

合成临床数据在医疗AI发展中日益重要,因真实电子健康记录(EHR)存在严格隐私限制、罕见病标注数据稀缺及观察性数据系统性偏差。尽管大语言模型(LLMs)能生成流畅临床文本,但确保合成数据兼具真实性与临床意义仍具挑战。我们提出DualAlign框架,通过双重对齐提升统计保真度与临床合理性:(1) 统计对齐,基于患者人口学特征和风险因素进行生成条件控制;(2) 语义对齐,引入真实症状演变轨迹引导内容生成。以阿尔茨海默病(AD)为例,DualAlign生成的上下文相关症状级句子更贴近真实临床文档。使用双源数据(DualAlign生成+人工标注)微调LLaMA 3.1-8B模型,在性能上显著优于仅用真实数据或无引导合成基线训练的模型。尽管未完全捕捉纵向复杂性,该方法为生成临床可信、隐私安全的合成数据提供实用路径,支持低资源临床文本分析。

原文摘要 · Abstract (English)

Synthetic clinical data are increasingly important for advancing AI in healthcare, given strict privacy constraints on real-world EHRs, limited availability of annotated rare-condition data, and systemic biases in observational datasets. While large language models (LLMs) can generate fluent clinical text, producing synthetic data that is both realistic and clinically meaningful remains challenging. We introduce DualAlign, a framework that enhances statistical fidelity and clinical plausibility through dual alignment: (1) statistical alignment, which conditions generation on patient demographics and risk factors; and (2) semantic alignment, which incorporates real-world symptom trajectories to guide content generation. Using Alzheimer's disease (AD) as a case study, DualAlign produces context-grounded symptom-level sentences that better reflect real-world clinical documentation. Fine-tuning an LLaMA 3.1-8B model with a combination of DualAlign-generated and human-annotated data yields substantial performance gains over models trained on gold data alone or unguided synthetic baselines. While DualAlign does not fully capture longitudinal complexity, it offers a practical approach for generating clinically grounded, privacy-preserving synthetic data to support low-resource clinical text analysis.

合成数据医疗AILLM隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。