arXiv:2607.03356eess.AS2026-07

构建心肺咳声谱图基准,评估模型对医疗音频的视觉理解能力

CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds

论文配图:CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds
图 1 · 摘自论文原文
  • 基于7个医学音频数据集生成梅尔频谱图,构建多模态问答基准
  • 9个顶尖视觉与通用模型最高准确率仅51.2%,难以捕捉细粒度声学特征
  • 揭示现有模型缺乏医疗声学图像理解能力,适合医疗多模态研究者参考

呼吸音、心脏听诊音和咳嗽音频蕴含丰富的诊断信息,但现有基准无法评估其频谱图表示下的多模态推理能力。我们提出CaReCoS,一个将临床问题与来自七个医学音频数据集的梅尔频谱图配对的基准。评估9个前沿视觉与全能模型发现,所有模型在频谱图中编码的细粒度声学特征上表现不佳:无一能可靠结合视觉模式识别与医学知识,最高准确率为51.2%,凸显了在医学声音可视化数据上训练的必要性。

原文摘要 · Abstract (English)

Medical acoustic signals such as respiratory sounds, cardiac auscultations, and cough audio carry rich diagnostic information, yet no existing benchmark evaluates multimodal reasoning over their spectrogram representations. We address both gaps with CaReCoS, a benchmark pairing clinically grounded questions with mel-spectrogram images derived from seven medical audio datasets. Evaluating 9 state-of-the-art vision and omni models, we find that all struggle with fine-grained acoustic features encoded in spectrograms: no model reliably combines visual pattern recognition with medical knowledge, achieving a maximum accuracy of 51.2%, underscoring the need for training on medical sound visualizations.

医疗音频频谱图多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。