arXiv:2504.18004eess.AScs.SD2025-04中稿 · IEEE EMBC 2025被引 2

评估现成音频基础模型在心肺音分析中的实用效果

Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis

  • 对比多个现成音频模型在四类心肺音任务上的表现
  • 噪声数据下性能差,清洁数据上达到当前最优水平
  • 通用音频模型优于专用呼吸音模型,适用性更广

预训练的深度学习模型(即基础模型)已成为自然语言处理和图像领域的重要组件。这一趋势已延伸至心肺音分析领域,其作为即插即用特征提取器已展现出有效性。然而,相关评估基准有限,导致与当前最优(SOTA)结果不兼容,阻碍了对其实际效能的验证。本研究通过对比多种现成音频基础模型在四类心肺音任务上的表现,评估其实际效用,并与SOTA微调结果进行比较。实验表明,在噪声数据下的两个任务中模型表现不佳,但在其余两个清洁数据任务上达到了SOTA性能。此外,通用音频模型的表现优于专用呼吸音模型,凸显其更广泛的应用潜力。研究还发布了代码,为未来心肺音基础模型的开发与应用提供支持。

原文摘要 · Abstract (English)

Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound models, which have demonstrated effectiveness as off-the-shelf feature extractors. However, their evaluation benchmarking has been limited, resulting in incompatibility with state-of-the-art (SOTA) performance, thus hindering proof of their effectiveness. This study investigates the practical effectiveness of off-the-shelf audio foundation models by comparing their performance across four respiratory and heart sound tasks with SOTA fine-tuning results. Experiments show that models struggled on two tasks with noisy data but achieved SOTA performance on the other tasks with clean data. Moreover, general-purpose audio models outperformed a respiratory sound model, highlighting their broader applicability. With gained insights and the released code, we contribute to future research on developing and leveraging foundation models for respiratory and heart sounds.

音频模型心肺音基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。