让医疗视觉模型像医生一样反复观察思考,提升诊断准确率。
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
- 设计迭代推理框架,模拟医生多轮观察与判断过程。
- 在16K医学问答数据上训练,显著优于现有模型。
- 适合医学AI研究者与临床辅助系统开发者。
医疗视觉语言模型(VLMs)在图文理解方面表现优异,但通常依赖单次推理,忽视局部视觉线索。在临床实践中,专家会反复扫描、聚焦并修正关注区域以做出最终诊断。为缩小机器与人类感知的差距,我们提出ViTAR,一种模仿专家迭代推理过程的新框架,采用“思考-行动-再思考-回答”的认知链。ViTAR将医学图像视为可交互对象,支持多步视觉推理。为此,我们构建了包含1000个交互示例的高质量指令数据集,刻画专家级诊断行为;同时整理了16000个用于细粒度诊断的视觉问答训练数据。采用两阶段训练策略:先通过监督微调引导认知轨迹,再通过强化学习优化决策。大量实验表明,ViTAR显著优于现有先进模型。注意力分析显示,从“思考”到“再思考”阶段,模型逐渐将视觉定位锚定于临床关键区域,并在推理过程中持续保持对视觉标记的关注,揭示其性能提升的机制。这些发现表明,在VLM中嵌入专家式迭代思维链,可提升医疗AI的性能与可信度。
原文摘要 · Abstract (English)
Medical vision-language models (VLMs) excel at image-text understanding but typically rely on a single-pass reasoning that neglects localized visual cues. In clinical practice, however, human experts iteratively scan, focus, and refine the regions of interest before reaching a final diagnosis. To narrow this machine-human perception gap, we introduce ViTAR, a novel VLM framework that emulates the iterative reasoning process of human experts through a cognitive chain of "think-act-rethink-answer". ViTAR treats medical images as interactive objects, enabling models to engage multi-step visual reasoning. To support this approach, we curate a high-quality instruction dataset comprising 1K interactive examples that encode expert-like diagnostic behaviors. In addition, a 16K visual question answering training data has been curated towards fine-grained visual diagnosis. We introduce a two-stage training strategy that begins with supervised fine-tuning to guide cognitive trajectories, followed by the reinforcement learning to optimize decision-making. Extensive evaluations demonstrate that ViTAR outperforms strong state-of-the-art models. Visual attention analysis reveals that from the "think" to "rethink" rounds, ViTAR increasingly anchors visual grounding to clinically critical regions and maintains high attention allocation to visual tokens during reasoning, providing mechanistic insight into its improved performance. These findings demonstrate that embedding expert-style iterative thinking chains into VLMs enhances both performance and trustworthiness of medical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。