arXiv:2607.22293cs.CV2026-07

提升医学影像理解可靠性,从底层视觉感知入手。

RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

论文配图:RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding
图 1 · 摘自论文原文
  • 设计双2D/3D编码器架构,保留原始影像空间结构。
  • 在113万样本的Perception-Bench上全面超越现有模型。
  • 适合追求临床诊断可靠性的医疗AI研究者使用。

医学多模态大语言模型(MLLMs)被寄望于完成复杂图像理解任务,但其可靠性常因视觉解读错误而受损。为系统追踪这些失败,我们从高层临床任务逐层下探至基础视觉感知。为此提出Perception-Bench,一个包含113万样本的大规模基准,评估MLLMs在六个维度的表现:属性判断、空间定位、空间理解、疾病预测、异常检测和报告生成,涵盖2D与3D放射影像。分析发现,现有模型连最基础的病灶位置、大小、密度等属性都无法准确捕捉。这种无法将临床输出锚定在原始视觉证据上的问题,揭示了诊断不可靠的根本瓶颈在于低层视觉感知。受此启发,我们提出RadSight,一种基于双2D/3D编码器架构的感知驱动型MLLM,旨在保持原生成像空间结构。该模型将医学图像理解建模为四个阶段的渐进过程:视觉-语言对齐、细粒度视觉感知、临床诊断与诊断解释。模型在837万条感知导向语料上通过渐进式课程学习进行训练。在Perception-Bench上,RadSight在全部六项评估维度上均显著优于现有模型,尤其在空间定位与临床诊断方面提升明显。同时,在公开的2D与3D医学基准上也展现出持续改进,进一步证明稳健的底层视觉感知是可靠临床理解的关键基础。代码与模型将公开。

原文摘要 · Abstract (English)

Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often compromised by frequent errors in visual interpretation. To systematically trace these failures, we traverse the hierarchy from high-level clinical tasks down to fundamental visual perception. We therefore introduce Perception-Bench, a large-scale benchmark comprising 1.13 million samples that assesses medical MLLMs across six dimensions: attribute judgment, spatial grounding, spatial understanding, disease prediction, anomaly detection, and report generation, spanning both 2D and 3D radiology images. Our analysis on Perception-Bench reveals that existing MLLMs lack the ability to capture even the most basic lesion attributes, such as location, size, and density. This inability to ground clinical outputs in primary visual evidence reveals that the models' diagnostic unreliability is rooted in a critical but overlooked bottleneck in low-level visual perception. Motivated by this, we propose RadSight, a perception-driven MLLM built upon a dual 2D/3D encoder architecture that preserves native imaging spatial structures. RadSight formulates medical image understanding as a four-stage progressive process: visual-language alignment, fine-grained visual perception, clinical diagnosis, and diagnostic interpretation. The model is trained on an 8.37 million perception-oriented corpus using progressive curriculum learning. On Perception-Bench, RadSight consistently outperforms existing MLLMs across all six evaluation dimensions, with particularly strong gains in spatial grounding and clinical diagnosis. It also achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable clinical understanding. Code and model will be publicly available.

医学影像多模态感知可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。