用扩散模型生成心脏音,评估其生理合理性与听觉真实性。
Diffusion-Based Heart Sound Generation: Evaluation with Physiological Signal Metrics, Classifiers, and Expert Listening

- 在对数梅尔域构建条件扩散模型生成心音信号。
- 生成信号保留正常/异常分类结构,但异常特征较弱。
- 专家听辨显示生成音类似真实心音,但异常敏感度不足。
公开的心音数据集规模小、病理多样性不足,限制了听诊训练和自动化分类器的泛化能力。本文在对数梅尔域开发了一种类别条件扩散模型用于心音生成,并通过三类方法评估合成质量:(i)基于生理特性的合理性指标,(ii)下游标签一致性测试,(iii)专家听辨。实验基于PhysioNet/Computing in Cardiology Challenge 2016数据集(3240条记录),经预处理后获得16,749个非重叠4秒片段,转换为标准化的1×128×128对数梅尔表示,训练带有无分类器引导的2D U-Net去噪器。重建波形使用三个轻量级指标量化信号层面合理性:包络自相关节律得分、振幅爆炸得分及主导周期滞后。结果显示,生成片段具有相似的主导周期时长,但包络周期性降低,瞬态爆发性增强。下游评估中,ResNet-50分类器在真实测试集上达92.24%准确率,在平衡类别的合成批次上达82.8%,表明生成信号仍保留对分类有意义的结构。初步专家听辨研究(60段,两名临床医生)显示多数合成片段被判定为心音样,但真实与合成片段的异常敏感度均较低。整体结果为基于扩散的心音生成提供了实用基准,同时指出了保留异常声学线索与减少重建伪影的挑战。
原文摘要 · Abstract (English)
Publicly available phonocardiogram (PCG) datasets remain limited in size and pathological diversity, constraining both auscultation training and the generalisation of automated heart-sound classifiers. A class-conditional diffusion model for PCG generation is developed in the log-mel domain and synthetic fidelity is assessed using complementary (i) physiology-inspired plausibility metrics, (ii) downstream label-consistency evaluation, and (iii) expert listening. Experiments use the Phy-sioNet/Computing in Cardiology Challenge 2016 dataset (3240 recordings) with recording-level splits. After preprocessing and quality control, 16,749 non-overlapping 4 s clips are mapped to a normalised 1 x 128 x 128 log-mel representation to train a conditional 2D U-Net denoiser with classifier-free guidance. Signal-level plausibility is quantified on reconstructed waveforms using three lightweight metrics: an envelope-autocorrelation rhythm score, an amplitude-based explosion score, and the dominant cycle lag. Synthetic clips preserve similar dominant cycle durations but exhibit reduced envelope periodicity and increased transient burstiness relative to real clips. For downstream evaluation, a ResNet-50 classifier achieves 92.24% accuracy on the held-out real test set and 82.8% accuracy on class-balanced synthetic batches, indicating that generated signals retain discriminative structure relevant to normal/abnormal classification. In a pilot expert listening study (60 clips, two clinicians), most synthetic clips are judged as heart-sound-like, while abnormality sensitivity is low for both real and synthetic 4 s excerpts. Overall, the results provide a practical baseline for diffusion-based PCG generation while highlighting remaining challenges in retaining abnormal acoustic cues and reducing reconstruction-induced artefacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。