arXiv:2504.17613cs.LG2025-04KDD被引 16

让生成的病历数据直接提升临床模型效果,不只像真病历。

TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series Generation

  • 用影响函数评估合成数据对下游任务的贡献,引导生成过程
  • 在6个公开数据集上,AUPRC提升20.4%,AUROC提升18.4%
  • 适合需要提升罕见病预测性能的医疗AI研究者

合成电子健康记录(EHR)时间序列生成对推动临床机器学习至关重要,可缓解数据稀缺问题。然而,现有方法多聚焦于复制真实数据的统计分布和时序依赖性。我们指出,仅保证与真实数据一致,并不能确保下游模型性能提升,因为常见模式可能主导,忽略重要但罕见的疾病。因此,需生成能提升特定临床模型表现的合成数据。为此,我们提出TarDiff,一种目标导向的扩散框架,将任务相关影响引导融入生成过程。不同于传统方法模仿训练数据分布,TarDiff通过影响函数量化合成样本对降低任务损失的预期贡献,并将该梯度嵌入反向扩散过程,从而生成更优效的数据。在六个公开EHR数据集上评估显示,其性能达到新高,AUPRC最高提升20.4%,AUROC最高提升18.4%。结果表明,TarDiff不仅保持时序保真度,还显著提升下游模型性能,为医疗数据分析中的数据稀缺与类别不平衡问题提供有效解决方案。

原文摘要 · Abstract (English)

Synthetic Electronic Health Record (EHR) time-series generation is crucial for advancing clinical machine learning models, as it helps address data scarcity by providing more training data. However, most existing approaches focus primarily on replicating statistical distributions and temporal dependencies of real-world data. We argue that fidelity to observed data alone does not guarantee better model performance, as common patterns may dominate, limiting the representation of rare but important conditions. This highlights the need for generate synthetic samples to improve performance of specific clinical models to fulfill their target outcomes. To address this, we propose TarDiff, a novel target-oriented diffusion framework that integrates task-specific influence guidance into the synthetic data generation process. Unlike conventional approaches that mimic training data distributions, TarDiff optimizes synthetic samples by quantifying their expected contribution to improving downstream model performance through influence functions. Specifically, we measure the reduction in task-specific loss induced by synthetic samples and embed this influence gradient into the reverse diffusion process, thereby steering the generation towards utility-optimized data. Evaluated on six publicly available EHR datasets, TarDiff achieves state-of-the-art performance, outperforming existing methods by up to 20.4% in AUPRC and 18.4% in AUROC. Our results demonstrate that TarDiff not only preserves temporal fidelity but also enhances downstream model performance, offering a robust solution to data scarcity and class imbalance in healthcare analytics.

医疗AI数据生成扩散模型时序数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。