用音频+大模型实现心肺音自动诊断,回答更像医生。
CaReAQA: A Cardiac and Respiratory Audio Question Answering Model for Open-Ended Diagnostic Reasoning
- 用基础音频模型+大语言模型做心肺音诊断推理
- 开放问答任务准确率达86.2%,闭合任务平均56.9%
- 适合临床辅助诊断系统研究者参考
心肺音等医学音频在临床诊断中至关重要。然而,传统方法依赖人工特征或需大量标注数据的监督深度学习模型,限制了其可扩展性与应用范围。为此,我们提出CaReAQA,一种融合基础音频模型与大语言模型推理能力的音视频-语言模型,可生成具有临床意义的开放式诊断回答。同时,我们构建了CaReSound,一个包含标注医学音频、元数据及配对问答示例的基准数据集,旨在推动诊断推理研究。评估显示,CaReAQA在开放式诊断推理任务中达到86.2%准确率,优于基线模型;在未见数据集上的闭合式分类任务平均准确率达56.9%。结果表明,音频-语言融合与推理能力可显著提升医疗诊断效率,助力临床决策支持系统发展。
原文摘要 · Abstract (English)
Medical audio signals, such as heart and lung sounds, play a crucial role in clinical diagnosis. However, analyzing these signals remains challenging: traditional methods rely on handcrafted features or supervised deep learning models that demand extensive labeled datasets, limiting their scalability and applicability. To address these issues, we propose CaReAQA, an audio-language model that integrates a foundation audio model with the reasoning capabilities of large language models, enabling clinically relevant, open-ended diagnostic responses. Alongside CaReAQA, we introduce CaReSound, a benchmark dataset of annotated medical audio recordings enriched with metadata and paired question-answer examples, intended to drive progress in diagnostic reasoning research. Evaluation results show that CaReAQA achieves 86.2% accuracy on open-ended diagnostic reasoning tasks, outperforming baseline models. It also generalizes well to closed-ended classification tasks, achieving an average accuracy of 56.9% on unseen datasets. Our findings show how audio-language integration and reasoning advances medical diagnostics, enabling efficient AI systems for clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。