arXiv:2509.23454eess.AScs.AI2025-09

融合频谱与波形信息,提升心音信号分类精度与鲁棒性

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

  • 用混合视觉变换器与一维卷积网络,同时处理心音频谱图和原始波形
  • 在PhysioNet数据集上达0.8608的ROC-AUC,优于单一模态模型
  • 对跨域数据表现更稳定,适合临床场景中模型泛化需求

心音信号(PCG)具有固有的周期性,在频谱(音调)和时间两个维度中包含诊断信息。传统二维频谱图虽能提取丰富频谱特征,但会损失相位信息和时间精度。本文提出AudioFuse,通过联合学习频谱图与原始波形的互补表示来分类心音信号。为降低融合模型过拟合风险,采用定制的宽而浅的视觉变压器(ViT)处理频谱图,配合浅层一维卷积神经网络处理原始波形。在PhysioNet 2016数据集上,从零训练即达到0.8608的ROC-AUC,优于仅使用频谱图(0.8066)或波形(0.8223)的基线模型。此外,在更具挑战性的PASCAL数据集上,其保持0.7181的ROC-AUC,而频谱基线模型性能骤降至0.4873。结果表明,融合互补表征可提供强归纳偏置,实现无需大规模预训练的高效、泛化性强分类器。

原文摘要 · Abstract (English)

Biomedical audio signals, such as phonocardiograms (PCG), are inherently rhythmic and contain diagnostic information in both their spectral (tonal) and temporal domains. Standard 2D spectrograms provide rich spectral features but compromise the phase information and temporal precision of the 1D waveform. We propose AudioFuse, an architecture that simultaneously learns from both complementary representations to classify PCGs. To mitigate the overfitting risk common in fusion models, we integrate a custom, wide-and-shallow Vision Transformer (ViT) for spectrograms with a shallow 1D CNN for raw waveforms. On the PhysioNet 2016 dataset, AudioFuse achieves a state-of-the-art competitive ROC-AUC of 0.8608 when trained from scratch, outperforming its spectrogram (0.8066) and waveform (0.8223) baselines. Moreover, it demonstrates superior robustness to domain shift on the challenging PASCAL dataset, maintaining an ROC-AUC of 0.7181 while the spectrogram baseline collapses (0.4873). Fusing complementary representations thus provides a strong inductive bias, enabling the creation of efficient, generalizable classifiers without requiring large-scale pre-training.

心音分析多模态学习音频分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。