arXiv:2505.20745cs.SDcs.LG2025-05被引 3

用基础模型表征提升心音听诊心率估计精度

Foundation Model Hidden Representations for Heart Rate Estimation from Auscultation

  • 分析六种自监督音频模型在心音数据中的层次表征
  • 自研CLAP模型编码器实现更低的平均绝对误差
  • 适合语音信号处理与医疗健康交叉研究者

听诊,尤其是心音检测,是一种提供关键生命体征信息的无创技术。近年来,自监督声学表征基础模型(FMs)被提出以揭示基于声音的生命体征信息。然而,目前对这些预训练模型中听诊信息编码程度的研究仍较少。本文使用公开的心电图(PCG)数据集和心率(HR)估计模型,对六种声学表征基础模型(HuBERT、wav2vec2、wavLM、Whisper、对比语言-音频预训练模型CLAP及自研CLAP模型)进行了逐层分析。同时,实现了Nie等人(2024)的基线方法(依赖声学特征),结果表明,预训练基础模型的表征向量整体性能可媲美基线。值得注意的是,使用自研CLAP模型音频编码器的表征,在不同训练/验证/测试划分下,均优于基线,取得了更低的平均绝对误差(MAE),尽管存在领域不匹配问题。

原文摘要 · Abstract (English)

Auscultation, particularly heart sound, is a non-invasive technique that provides essential vital sign information. Recently, self-supervised acoustic representation foundation models (FMs) have been proposed to offer insights into acoustics-based vital signs. However, there has been little exploration of the extent to which auscultation is encoded in these pre-trained FM representations. In this work, using a publicly available phonocardiogram (PCG) dataset and a heart rate (HR) estimation model, we conduct a layer-wise investigation of six acoustic representation FMs: HuBERT, wav2vec2, wavLM, Whisper, Contrastive Language-Audio Pretraining (CLAP), and an in-house CLAP model. Additionally, we implement the baseline method from Nie et al., 2024 (which relies on acoustic features) and show that overall, representation vectors from pre-trained foundation models (FMs) offer comparable performance to the baseline. Notably, HR estimation using the representations from the audio encoder of the in-house CLAP model outperforms the results obtained from the baseline, achieving a lower mean absolute error (MAE) across various train/validation/test splits despite the domain mismatch.

心率估计基础模型音频表征医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。