用医学语言模型教音频模型理解听诊意义,提升诊断准确率。
Language Models as Semantic Teachers: Post-Training Alignment for Medical Audio Understanding
- 用语言模型当语义教师,对齐音频编码器与临床含义。
- 在18项任务上平均AUROC从0.68提升至0.79,新冠咳嗽检测达0.89。
- 适合医疗语音分析、可解释性诊断系统研究者使用。
预训练音频模型擅长识别听诊音中的声学模式,但难以理解其临床意义,限制了其在诊断任务中的应用与表现。为弥合这一差距,我们提出AcuLa(Audio-Clinical Understanding via Language Alignment),一种轻量级后训练框架,通过将任意音频编码器与医学语言模型对齐,赋予其语义理解能力。为实现大规模对齐,我们利用现成的大语言模型,将现有音频记录中丰富的结构化元数据转换为连贯的临床报告,构建大规模数据集。对齐策略结合表示层对比目标与自监督建模,确保模型学习临床语义的同时保留精细的时间线索。AcuLa在来自10个不同数据集的18个心肺任务中达到最先进水平,分类基准平均AUROC从0.68提升至0.79;在最具挑战性的新冠咳嗽检测任务中,AUROC从0.55跃升至0.89。本工作证明,音频-语言对齐可将纯声学模型转化为具有临床意识的诊断工具,确立了提升基于音频的生理理解的新范式。
原文摘要 · Abstract (English)
Pre-trained audio models excel at detecting acoustic patterns in auscultation sounds but often fail to grasp their clinical significance, limiting their use and performance in diagnostic tasks. To bridge this gap, we introduce AcuLa (Audio-Clinical Understanding via Language Alignment), a lightweight post-training framework that instills semantic understanding into any audio encoder by aligning it with a medical language model, which acts as a "semantic teacher." To enable alignment at scale, we construct a large-scale dataset by leveraging off-the-shelf large language models to translate the rich, structured metadata accompanying existing audio recordings into coherent clinical reports. Our alignment strategy combines a representation-level contrastive objective with a self-supervised modeling, ensuring that the model learns clinical semantics while preserving fine-grained temporal cues. AcuLa achieves state-of-the-art results across 18 diverse cardio-respiratory tasks from 10 different datasets, improving the mean AUROC on classification benchmarks from 0.68 to 0.79 and, on the most challenging COVID-19 cough detection task, boosting the AUROC from 0.55 to 0.89. Our work demonstrates that this audio-language alignment transforms purely acoustic models into clinically-aware diagnostic tools, establishing a novel paradigm for enhancing physiological understanding in audio-based health monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。