arXiv:2509.11606cs.SDcs.LG2025-09被引 1

用生成模型扩充心音数据,提升多模态心音分类准确率

Scaling to Multimodal and Multichannel Heart Sound Classification with Synthetic and Augmented Biosignals

  • 结合传统信号处理与扩散模型生成新心音数据
  • 在多个数据集上达92%以上准确率,最优结果超现有方法
  • 适合医疗AI研究者及可穿戴设备开发人员参考

心血管疾病是全球主要死因,每年约1790万人因此死亡。早期检测至关重要,亟需准确且低成本的筛查手段。深度学习已用于分析同步的心音图(PCG)与心电图(ECG)信号,以及多通道心音(mPCG)。但因同步多通道数据稀缺,先进架构难以充分发挥。本研究结合传统信号处理与去噪扩散模型WaveGrad和DiffWave,构建增强数据集,以微调基于Wav2Vec 2.0的心音分类器。在CinC 2016单通道PCG数据集上,准确率、未加权平均召回率(UAR)、敏感性、特异性及马修相关系数(MCC)分别达92.48%、93.05%、93.63%、92.48%、94.93%和0.8283。使用CinC训练A数据集的同步PCG+ECG信号,各项指标为93.14%、92.21%、94.35%、90.10%、95.12%和0.8380。在可穿戴背心采集的mPCG数据集上,准确率为77.13%,UAR为74.25%,敏感性为86.47%,特异性为62.04%,MCC为0.5082。结果表明,在增强数据支持下,基于Transformer的模型能有效实现心血管病检测。

原文摘要 · Abstract (English)

Cardiovascular diseases (CVDs) are the leading cause of death worldwide, accounting for approximately 17.9 million deaths each year. Early detection is critical, creating a demand for accurate and inexpensive pre-screening methods. Deep learning has recently been applied to classify abnormal heart sounds indicative of CVDs using synchronised phonocardiogram (PCG) and electrocardiogram (ECG) signals, as well as multichannel PCG (mPCG). However, state-of-the-art architectures remain underutilised due to the limited availability of synchronised and multichannel datasets. Augmented datasets and pre-trained models provide a pathway to overcome these limitations, enabling transformer-based architectures to be trained effectively. This work combines traditional signal processing with denoising diffusion models, WaveGrad and DiffWave, to create an augmented dataset to fine-tune a Wav2Vec 2.0-based classifier on multimodal and multichannel heart sound datasets. The approach achieves state-of-the-art performance. On the Computing in Cardiology (CinC) 2016 dataset of single channel PCG, accuracy, unweighted average recall (UAR), sensitivity, specificity and Matthew's correlation coefficient (MCC) reach 92.48%, 93.05%, 93.63%, 92.48%, 94.93% and 0.8283, respectively. Using the synchronised PCG and ECG signals of the training-a dataset from CinC, 93.14%, 92.21%, 94.35%, 90.10%, 95.12% and 0.8380 are achieved for accuracy, UAR, sensitivity, specificity and MCC, respectively. Using a wearable vest dataset consisting of mPCG data, the model achieves 77.13% accuracy, 74.25% UAR, 86.47% sensitivity, 62.04% specificity, and 0.5082 MCC. These results demonstrate the effectiveness of transformer-based models for CVD detection when supported by augmented datasets, highlighting their potential to advance multimodal and multichannel heart sound classification.

心音分类生成模型多模态医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。