将多部位听诊信号对齐大语言模型,实现患者级综合诊断
Patient-Level Multimodal Question Answering from Multi-Site Auscultation Recordings
- 用门控交叉注意力对齐听诊信号与冻结LLM的嵌入空间
- 在CaReSound上达0.865 F1-macro和0.952 BERTScore
- 轻量级专用编码器可媲美大型音频-语言模型
听诊是重要的诊断工具,但其应用常受限于主观判断。尽管通用音频-语言模型(ALMs)在一般领域表现优异,却难以处理生理信号的细微特征。本文提出一种框架,通过门控交叉注意力将多部位听诊记录直接对齐至冻结的大语言模型(LLM)嵌入空间。利用LLM的潜在世界知识,该方法超越单一分类,实现患者级的综合性评估。在CaReSound基准上,模型达到0.865 F1-macro和0.952 BERTScore的领先性能。实验表明,轻量级、领域特定的编码器可媲美大规模ALMs,且多部位聚合提供了空间冗余,缓解了时间截断问题。该方法为连接信号处理与临床评估提供了一条可扩展路径。
原文摘要 · Abstract (English)
Auscultation is a vital diagnostic tool, yet its utility is often limited by subjective interpretation. While general-purpose Audio-Language Models (ALMs) excel in general domains, they struggle with the nuances of physiological signals. We propose a framework that aligns multi-site auscultation recordings directly with a frozen Large Language Model (LLM) embedding space via gated cross-attention. By leveraging the LLM's latent world knowledge, our approach moves beyond isolated classification toward holistic, patient-level assessment. On the CaReSound benchmark, our model achieves a state-of-the-art 0.865 F1-macro and 0.952 BERTScore. We demonstrate that lightweight, domain-specific encoders rival large-scale ALMs and that multi-site aggregation provides spatial redundancy that mitigates temporal truncation. This alignment of medical acoustics with text foundations offers a scalable path for bridging signal processing and clinical assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。