arXiv:2501.13884eess.AScs.AI2025-01被引 11

用音频大模型分析心跳杂音,能识别11种医生标注特征

Exploring Finetuned Audio-LLM on Heart Murmur Features

  • 微调音频大模型Qwen2-Audio处理心音信号
  • 在11项特征中8项超越现有方法,长尾特征也表现良好
  • 适合心血管诊断辅助,尤其对数据少的罕见特征有效

语音大模型在人声、音乐和环境声识别上表现优异,但对生物医学声音(如心音)的潜力尚未充分探索。本研究聚焦于通过心音图(PCG)诊断心血管疾病,现有深度神经网络仅能区分健康与异常,无法预测杂音的时间、分级、粗糙度、音高和性质等关键特征。我们对Qwen2-Audio模型在PhysioNet CirCor DigiScope PCG数据集上进行微调,并评估其对11项专家标注杂音特征的分类性能。同时,采用音频表示模型SSAMBA的预处理分割算法以提升抗噪性和泛化能力。结果表明,该模型在11项特征中有8项优于当前最优方法,其余3项持平;且成功识别了训练数据稀少的长尾特征,而此前所有方法均未能实现。这证明音频大模型可作为心内科医生的诊断助手,显著提升心脏病识别能力。

原文摘要 · Abstract (English)

Large language models (LLMs) for audio have excelled in recognizing and analyzing human speech, music, and environmental sounds. However, their potential for understanding other types of sounds, particularly biomedical sounds, remains largely underexplored despite significant scientific interest. In this study, we focus on diagnosing cardiovascular diseases using phonocardiograms, i.e., heart sounds. Most existing deep neural network (DNN) paradigms are restricted to heart murmur classification (healthy vs unhealthy) and do not predict other acoustic features of the murmur such as timing, grading, harshness, pitch, and quality, which are important in helping physicians diagnose the underlying heart conditions. We propose to finetune an audio LLM, Qwen2-Audio, on the PhysioNet CirCor DigiScope phonocardiogram (PCG) dataset and evaluate its performance in classifying 11 expert-labeled murmur features. Additionally, we aim to achieve more noise-robust and generalizable system by exploring a preprocessing segmentation algorithm using an audio representation model, SSAMBA. Our results indicate that the LLM-based model outperforms state-of-the-art methods in 8 of the 11 features and performs comparably in the remaining 3. Moreover, the LLM successfully classifies long-tail murmur features with limited training data, a task that all previous methods have failed to classify. These findings underscore the potential of audio LLMs as assistants to human cardiologists in enhancing heart disease diagnosis.

音频大模型心音分析医疗诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。