arXiv:2503.06096cs.LG2025-03被引 6

用注意力机制生成真实医疗数据,提升慢性肾病生存分析模型校准精度。

Attention-Based Synthetic Data Generation for Calibration-Enhanced Survival Analysis: A Case Study for Chronic Kidney Disease Using Electronic Health Records

  • 基于注意力机制生成合成电子病历,保留关键临床指标
  • 整体校准误差降低15%,10个亚组平均降低9%
  • 适合需要数据隐私保护和精准建模的医疗研究者

真实世界医疗数据受限于严格隐私法规和数据不平衡,阻碍了研究与临床应用发展。合成数据虽具潜力,但现有方法难以保证真实性、可用性与校准性。本文提出掩码临床建模(MCM)框架,一种基于注意力机制的方法,可生成高保真合成数据集,保留危险比等关键临床信息,并增强生存模型校准能力。相比传统统计方法(如SMOTE)和机器学习模型(如VAEs),MCM支持独立数据合成与条件增广,满足多样化研究需求。在慢性肾病电子健康记录数据集上验证,MCM使全数据集校准损失降低15%;在10个临床分层子组中,平均校准损失降低9%,优于15种替代方法。该方法推动医疗模型精准化,提升稀缺医疗资源利用效率。

原文摘要 · Abstract (English)

Access to real-world healthcare data is limited by stringent privacy regulations and data imbalances, hindering advancements in research and clinical applications. Synthetic data presents a promising solution, yet existing methods often fail to ensure the realism, utility, and calibration essential for robust survival analysis. Here, we introduce Masked Clinical Modelling (MCM), an attention-based framework capable of generating high-fidelity synthetic datasets that preserve critical clinical insights, such as hazard ratios, while enhancing survival model calibration. Unlike traditional statistical methods like SMOTE and machine learning models such as VAEs, MCM supports both standalone dataset synthesis for reproducibility and conditional simulation for targeted augmentation, addressing diverse research needs. Validated on a chronic kidney disease electronic health records dataset, MCM reduced the general calibration loss over the entire dataset by 15%; and MCM reduced a mean calibration loss by 9% across 10 clinically stratified subgroups, outperforming 15 alternative methods. By bridging data accessibility with translational utility, MCM advances the precision of healthcare models, promoting more efficient use of scarce healthcare resources.

医疗生成生存分析合成数据注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。