arXiv:2510.05492cs.LGcs.AI2025-10

用频谱引导的扩散模型生成高保真心电图,兼顾临床真实与隐私安全。

High-Fidelity Synthetic ECG Generation via Mel-Spectrogram Informed Diffusion Training

  • 引入频谱域监督训练,提升心电图波形结构真实性。
  • 合成信号在10个指标上优于基线4-8%,跨导联相关误差降74%。
  • 支持个性化生成,适合低数据场景下的医疗AI训练。

心脏病学机器学习的发展因真实患者心电图(ECG)数据的隐私限制而受阻。尽管生成式AI提供潜在解决方案,但现有模型合成的ECG在可信度和临床实用性方面仍存在明显不足。本文针对当前生成式ECG方法的两大缺陷——形态保真度不足、无法生成个性化生理信号——提出两种创新:(1)基于条件扩散的结构状态空间模型(SSSD-ECG),采用新型训练范式MIDT-ECG(Mel-Spectrogram Informed Diffusion Training),通过时频域监督强化生理结构真实性;(2)多模态人口统计学条件输入,实现患者特定信号生成。我们在PTB-XL数据集上全面评估该方法,涵盖保真度、临床一致性、隐私保护及下游任务效用。结果表明,MIDT-ECG显著提升形态一致性,所有评估指标均优于基线4-8%,平均降低跨导联相关误差74%;人口统计学条件使信噪比与个性化程度提升。在低数据条件下,使用合成数据增强的分类器性能接近仅用真实数据训练的模型。证明该方法可在真实数据稀缺时,作为高保真、个性化、隐私安全的心电图替代品,推动生成式AI在医疗领域的负责任应用。

原文摘要 · Abstract (English)

The development of machine learning for cardiac care is severely hampered by privacy restrictions on sharing real patient electrocardiogram (ECG) data. Although generative AI offers a promising solution, the real-world use of existing model-synthesized ECGs is limited by persistent gaps in trustworthiness and clinical utility. In this work, we address two major shortcomings of current generative ECG methods: insufficient morphological fidelity and the inability to generate personalized, patient-specific physiological signals. To address these gaps, we build on a conditional diffusion-based Structured State Space Model (SSSD-ECG) with two principled innovations: (1) MIDT-ECG (Mel-Spectrogram Informed Diffusion Training), a novel training paradigm with time-frequency domain supervision to enforce physiological structural realism, and (2) multi-modal demographic conditioning to enable patient-specific synthesis. We comprehensively evaluate our approach on the PTB-XL dataset, assessing the synthesized ECG signals on fidelity, clinical coherence, privacy preservation, and downstream task utility. MIDT-ECG achieves substantial gains: it improves morphological coherence, preserves strong privacy guarantees with all metrics evaluated exceeding the baseline by 4-8%, and notably reduces the interlead correlation error by an average of 74%, while demographic conditioning enhances signal-to-noise ratio and personalization. In critical low-data regimes, a classifier trained on datasets supplemented with our synthetic ECGs achieves performance comparable to a classifier trained solely on real data. Together, we demonstrate that ECG synthesizers, trained with the proposed time-frequency structural regularization scheme, can serve as personalized, high-fidelity, privacy-preserving surrogates when real data are scarce, advancing the responsible use of generative AI in healthcare.

心电图生成扩散模型隐私保护医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。