arXiv:2607.12645cs.LG2026-07

解决电子病历生成中罕见病症低频问题,提升稀有群体数据真实性。

AdaPCLA: Adaptive Prior-Calibrated Logit Adjustment for Long-Tailed Longitudinal EHR Generation

论文配图:AdaPCLA: Adaptive Prior-Calibrated Logit Adjustment for Long-Tailed Longitudinal EHR Generation
图 1 · 摘自论文原文
  • 基于模拟退火内化数据分布先验,自适应调整生成逻辑。
  • 在MIMIC-III上尾部事件生成准确率提升114.2%,零样本跨人群适配F1增3.5%。
  • 无需再训练即可适应不同临床人群,适合医疗数据合成与隐私研究。

纵向电子病历的生成对隐私保护研究日益重要,但标准自回归模型常低估尾部事件(如罕见疾病、症状)的共现结构,导致生成数据对稀有亚群缺乏保真度。为此,我们提出AdaPCLA框架,通过数据分布感知的训练策略,使生成模型能自适应拟合并生成电子病历;该策略通过模拟退火训练内化数据知识参数。该方法还支持零样本分布控制,无需再训练即可适应多样化临床人群。理论分析揭示了基于标签的实证神经正切核(NTK)对稀有代码逻辑值更新的影响,并推导出退火速度与NTK条件对保留先验信号的边界。在真实世界数据上的实验表明,AdaPCLA在尾部合理性、下游任务效用和零样本控制方面均实现稳定提升;尤其在MIMIC-III上,尾部配对可见性(TailPairSeen)相比HALO提升114.2%,在MIMIC-IV上提升65.1%;在零样本跨人群适配上,比GPT类生成提升3.5% F1得分。

原文摘要 · Abstract (English)

Generative modeling of longitudinal Electronic Health Records is increasingly important for privacy-preserving research, yet standard autoregressive models tend to underrepresent the co-occurrence structure of tail events (i.e., diseases, symptoms), reducing the fidelity and faithfulness of generated data for rare subpopulations. To this end, we propose AdaPCLA framework, which enables generative models to adaptively fit and generate EHR data through a data distribution-aware training strategy; this is achieved by internalizing data knowledge parameters by simulated annealing training. It also supports training-free adaptation to a diverse clinical population for generation through zero-shot distribution control. Moreover, our theoretical analysis characterizes rare-code logit updates through the label-wise empirical NTK and derives a prior-internalization bound for how annealing speed and NTK conditioning affect retained prior signals. Experiments on real-world data show that AdaPCLA achieves consistent gains in tail plausibility, downstream utility, and zero-shot control; in particular, it improves TailPairSeen over HALO by 114.2% on MIMIC-III and 65.1% on MIMIC-IV, outperforms GPT-style generation by 3.5% F1 for zero-shot cross-population adaptation.

电子病历生成长尾分布零样本生成医疗数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。