arXiv:2510.02322eess.AScs.CL2025-10被引 1

用语音报告训练医学影像模型,让语音也能参与诊断分析。

SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis

  • 从语音报告中学习视觉语言表示,构建语音-影像对齐模型。
  • 零样本分类F1提升至0.705,恢复88%性能差距。
  • 无需文本即可实现精准检索,适合临床语音交互场景。

口语化沟通在临床工作中至关重要,例如放射科报告多通过口述生成。然而,现有医疗AI系统几乎仅依赖书面文本。本文探索直接从口头放射科报告中学习视觉-语言表示的可行性。我们构建了大规模语音报告数据集Speech-RATE,训练了SpeechCT-CLIP模型,将语音与3D CT图像映射到共享表示空间。尽管基础语音模型表现逊于文本模型,但通过从预训练的文本-图像CLIP模型进行知识蒸馏,成功将语义对齐能力迁移至语音,显著缩小性能差距。实验显示,零样本分类F1从0.623提升至0.705,恢复88%的性能差值,且推理阶段无需文本即可实现优异检索效果。结果表明语音是多模态预训练中可实用的替代文本方案,为临床语音驱动的辅助诊断工具开辟新路径。

原文摘要 · Abstract (English)

Spoken communication plays a central role in clinical workflows. In radiology, for example, most reports are created through dictation. Yet, nearly all medical AI systems rely exclusively on written text. In this work, we address this gap by exploring the feasibility of learning visual-language representations directly from spoken radiology reports. Specifically, we synthesize a large-scale dataset (Speech-RATE) of spoken radiology reports and train SpeechCT-CLIP, a contrastive model that aligns speech and 3D CT volumes in a shared representation space. While naive speech-based models underperform compared to text-trained counterparts, we show that knowledge distillation from a pretrained text-image CLIP model effectively transfers semantic alignment capabilities from text to speech, substantially narrowing this gap. Experiments demonstrate improved zero-shot classification F1 from 0.623 to 0.705, recovering 88% of the performance difference, and strong retrieval results without requiring text at inference. These findings highlight speech as a practical alternative to text in multimodal pretraining and open the door to voice-driven diagnostic support tools in clinical practice.

语音理解医学影像多模态知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。