用隐私保护方法生成真实临床笔记,效果接近原始数据。
Privacy-Preserving Generation of Clinical Narratives from Medical Terminologies
- 分离内容与形式,分段生成带医学术语的临床笔记。
- 合成笔记统计特征接近真实数据,下游模型性能相当。
- 无需标签分布假设,适合医疗等敏感领域应用。
在医疗等高风险领域,隐私问题严重限制了真实数据的使用。差分隐私(DP)合成数据提供了正式隐私保障的替代方案,但因领域专属性和长文本复杂性,临床笔记生成的实用价值仍难实现。我们提出Term2Note,一种在差分隐私约束下生成完整临床笔记的方法。通过结构化分离内容与形式,模型以医学术语为条件,分段生成笔记内容,并对术语与笔记分别施加独立的差分隐私约束,同时引入差分隐私质量最大化器优化输出。实验表明,Term2Note生成的合成笔记在统计特性上与真实临床笔记高度一致,基于这些合成数据训练的下游模型性能可媲美使用真实数据训练的模型。相比现有差分隐私文本生成基线,Term2Note显著提升生成质量与实用性,且不依赖标签分布假设,展现出作为真实临床笔记隐私保护替代方案的可行性。
原文摘要 · Abstract (English)
In high-stakes domains such as healthcare, privacy concerns severely limit the use of real-world training data. Differentially private (DP) synthetic data offers a promising alternative with formal privacy guarantees, but achieving strong utility remains challenging for clinical note generation due to domain specificity and long-form text complexity. We present Term2Note, a method for synthesising full-length clinical notes under DP constraints. By structurally separating content and form, Term2Note generates section-wise note content conditioned on medical terms, with terms and notes privatised under separate DP constraints, and applies a DP quality maximiser to improve outputs. Experiments demonstrate that Term2Note produces synthetic notes with statistical properties closely aligned with real clinical notes, and that downstream models trained on these notes achieve performance comparable to those trained on real clinical data. Compared to existing DP text generation baselines, Term2Note substantially improves both fidelity and utility, without relying on label distribution assumptions, highlighting its effectiveness as a practical privacy-preserving alternative to real clinical notes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。