用无标签数据训练心脏音频模型,少3倍标注数据仍超顶尖表现
SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation
- 双路径对比学习:同时处理波形和频谱图,融合多模态信息
- 仅需1/3标注数据,性能媲美全监督模型,跨任务泛化能力强
- 适合医疗音频分析、小样本学习及可解释性要求高的场景
心音图的自动化分析对心血管疾病早期诊断至关重要,但监督深度学习受限于专家标注数据稀缺。本文提出自监督双路径原型网络(SS-DPPN),一种从无标签数据中学习心脏音频表示与分类的基础模型。该框架采用基于对比学习的双路径结构,通过新颖的混合损失函数同步处理一维波形与二维频谱图。下游任务采用原型网络进行度量学习,提升模型敏感性并生成校准可靠的预测结果。SS-DPPN在四个心脏音频基准上达到当前最优性能,仅需三倍减少标注数据即可超越全监督模型,且所学表征成功泛化至肺音分类与心率估计任务。实验验证了其作为生理信号基础模型的鲁棒性、可靠性与可扩展性。
原文摘要 · Abstract (English)
The automated analysis of phonocardiograms is vital for the early diagnosis of cardiovascular disease, yet supervised deep learning is often constrained by the scarcity of expert-annotated data. In this paper, we propose the Self-Supervised Dual-Path Prototypical Network (SS-DPPN), a foundation model for cardiac audio representation and classification from unlabeled data. The framework introduces a dual-path contrastive learning based architecture that simultaneously processes 1D waveforms and 2D spectrograms using a novel hybrid loss. For the downstream task, a metric-learning approach using a Prototypical Network was used that enhances sensitivity and produces well-calibrated and trustworthy predictions. SS-DPPN achieves state-of-the-art performance on four cardiac audio benchmarks. The framework demonstrates exceptional data efficiency with a fully supervised model on three-fold reduction in labeled data. Finally, the learned representations generalize successfully across lung sound classification and heart rate estimation. Our experiments and findings validate SS-DPPN as a robust, reliable, and scalable foundation model for physiological signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。