arXiv:2608.26856cs.CVcs.AI2026-08中稿 · ECCV

让医学多模态模型从推理到像素精准定位,提升诊断可信度。

From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

论文配图:From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
图 1 · 摘自论文原文
  • 通过分割锚定推理池化,从模型内部提取与任务相关的语义证据。
  • 在基准测试中达到68.49% gIoU和70.47% cIoU,显著优于现有方法。
  • 适合需要可解释医学图像分析的临床研究与医疗AI开发人员。

尽管多模态大语言模型(MLLMs)在医学视觉问答(Med-VQA)上表现优异,但其依赖全局图像特征,缺乏像素级定位,限制了临床可信度。为弥合高层临床推理与空间定位之间的语义鸿沟,我们提出统一框架 extsc{MedREAL}(Medical Reasoning-driven Answering and Localization),实现语言推理与空间定位的无缝对齐。具体而言, extsc{MedREAL} 引入分割锚定推理池化(SARP),直接从 MLLM 隐状态中的 exttt{[SEG]} token 中提炼任务相关语义证据。此外,设计推理到视觉(R2V)融合机制,将这些推理感知特征注入分割流水线以实现精确掩码解码。为支持该范式,我们构建了包含13,824个专家验证样本的 MedRAVS-13K 数据集,覆盖四种不同成像模态。大量实验表明, extsc{MedREAL} 显著优于现有最先进方法,在基准评估中达到68.49% gIoU和70.47% cIoU。通过生成与文本诊断严格一致的证据掩码, extsc{MedREAL} 提供了一个强大且可解释的推理驱动医学图像分析框架。

原文摘要 · Abstract (English)

Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA), their reliance on global image features often lacks precise pixel-level grounding, thereby limiting clinical trustworthiness. To bridge the semantic gap between high-level clinical reasoning and spatial localization, we propose \textsc{\textsc{MedREAL}} (\textbf{Med}ical \textbf{RE}asoning-driven \textbf{A}nswering and \textbf{L}ocalization), a unified framework that seamlessly aligns linguistic reasoning with spatial grounding. Specifically, \textsc{MedREAL} introduces \textbf{S}eg \textbf{A}nchored \textbf{R}easoning \textbf{P}ooling (SARP) to distill task-relevant semantic evidence directly from \texttt{[SEG]} tokens within the MLLM's hidden states. Furthermore, a \textbf{R}easoning-to-\textbf{V}isual (R2V) fusion mechanism is proposed to effectively inject these reasoning-aware features into a segmentation pipeline for accurate mask decoding. To facilitate this paradigm, we construct MedRAVS-13K, a comprehensive dataset comprising 13,824 expertly validated samples across four diverse imaging modalities. Extensive experiments demonstrate that \textsc{MedREAL} significantly outperforms state-of-the-arts, achieving 68.49\% gIoU and 70.47\% cIoU on benchmark evaluations. By generating evidence masks that are strictly consistent with textual diagnoses, \textsc{MedREAL} provides a robust, interpretable framework for reasoning-driven medical image analysis.

医学多模态视觉问答分割定位可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。