arXiv:2607.24743cs.CVcs.AI2026-07

ClinFusion提升医学影像理解,让AI报告更精准可信。

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

论文配图:ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
图 1 · 摘自论文原文
  • 构建级联视觉编码器,统一处理2D/3D医学图像
  • 在24个基准中20项超越开源模型,16项胜过闭源大模型
  • 医生盲评显示报告质量最高,评估指标与专家判断最相关

多模态大语言模型在临床应用中潜力巨大,但其部署本质是视觉中心挑战:模型需从异构的二维和三维医学影像中吸收知识,评估方式须符合放射科医生的实际工作流程,并实现精确、细粒度且以事实为基础的评估。本文提出ClinFusion,一种面向整体医学理解的视觉中心型多模态大模型,系统解决上述问题。我们设计了一种组合式级联视觉编码器架构,包含级联空间感知局部融合算子,将多样化的2D与原生3D医学图像理解统一于一个融合编码器中。进一步提出基于视觉的评估框架,包括用于指令遵循评估的MedIF-Bench,以及基于感兴趣区域(RoI)的临床对齐且事实驱动的报告生成评估方法。实验表明,ClinFusion在涵盖视觉问答、报告生成和指令遵循的2D与3D多模态医学基准,以及文本医学任务上均达到新标杆表现,在24个基准中有20项优于主流开源医疗多模态模型(如Hulu-Med、Lingshu),在16个基准中优于强大闭源模型(如GPT-5.2、Gemini-3-Flash)。该模型还可通过代理工具调用增强检索增强与工具辅助的临床工作流。由认证放射科医生进行的盲评确认,ClinFusion生成报告排名最高,且其提出的RoI基评估指标与专家判断的相关性最强。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (\textit{e.g.}, Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.

医学AI多模态视觉理解生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。