让医生用语音指挥模型看片子,还能解释判断依据。
SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging
- 用语音替代文字控制医疗影像分析,适合手术等场景
- 构建新数据集,实现诊断结果的可解释推理
- 适合临床医生、医疗AI研究者关注
医学视觉语言模型在医疗图像描述和辅助诊断中展现出巨大潜力。然而,现有模型多依赖文本指令,在手术等场景中难以使用;且多数模型缺乏对异常判断的完整推理过程,影响临床可信度。为此,我们提出端到端语音驱动的医学VLM——SilVar-Med,首次实现语音交互式医疗影像分析。同时,构建了用于推理解释的新数据集,通过实验证明了语音驱动与推理解释结合的可行性。本工作有望推动更透明、交互性更强、临床可用的智能诊断系统发展。代码与数据集已公开于SiVar-Med。
原文摘要 · Abstract (English)
Medical Visual Language Models have shown great potential in various healthcare applications, including medical image captioning and diagnostic assistance. However, most existing models rely on text-based instructions, limiting their usability in real-world clinical environments especially in scenarios such as surgery, text-based interaction is often impractical for physicians. In addition, current medical image analysis models typically lack comprehensive reasoning behind their predictions, which reduces their reliability for clinical decision-making. Given that medical diagnosis errors can have life-changing consequences, there is a critical need for interpretable and rational medical assistance. To address these challenges, we introduce an end-to-end speech-driven medical VLM, SilVar-Med, a multimodal medical image assistant that integrates speech interaction with VLMs, pioneering the task of voice-based communication for medical image analysis. In addition, we focus on the interpretation of the reasoning behind each prediction of medical abnormalities with a proposed reasoning dataset. Through extensive experiments, we demonstrate a proof-of-concept study for reasoning-driven medical image interpretation with end-to-end speech interaction. We believe this work will advance the field of medical AI by fostering more transparent, interactive, and clinically viable diagnostic support systems. Our code and dataset are publicly available at SiVar-Med.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。