让AI像医生一样先聚焦病灶再诊断,提升超声图像问答准确率。
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming

- 设计可交互缩放的诊断流程,模拟医生聚焦病灶的思考方式。
- 通过不确定性奖励机制,使模型在模糊情况保持谨慎,清晰时更自信。
- 在肝、乳腺、甲状腺数据集上定位精度提升39.3%,适合临床辅助诊断场景。
视觉语言模型(VLM)在医学图像问答中进展显著,但在超声图像上的表现仍不理想。临床上,超声医师会主动聚焦病灶区域进行诊断,但因主观性导致判断存在差异。现有VLM未显式支持诊断前的交互式缩放,且通常将标注视为无偏真值,忽略了其固有的主观性和模糊性。本文提出一种新框架,模拟超声医师的认知流程。首先引入结构化的‘缩放-诊断’范式,实现病灶聚焦推理;其次,在组相对策略优化(GRPO)框架中,基于随机分组滚动推演设计不确定性感知奖励,以预测一致性作为模型置信度的代理。二者协同促使模型在明确病例中强化准确预测,模糊情况下保持审慎。在肝脏、乳腺和甲状腺数据集上的实验表明,该框架使病灶定位准确率提升39.3%,证明模型已学会主动靠近观察并作出诊断。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have significantly advanced medical visual question answering, yet their performance in ultrasound remains suboptimal. In clinical practice, sonographers explicitly focus on lesion regions to formulate reports, though diagnostic interpretations sometimes vary due to inherent subjectivity. However, existing VLMs are not explicitly structured to interactively zoom into lesions prior to diagnosis; moreover, they typically treat annotations as unbiased ground truths, failing to account for their inherent subjectivity and ambiguity. In this paper, we propose a framework specifically designed to consider the sonographer's cognitive workflow. We first introduce a structured Zoom-then-Diagnose paradigm, which replicates the interactive search process to enable lesion-focused reasoning. Furthermore, within the Group Relative Policy Optimization (GRPO) framework, we introduce an uncertainty-aware reward derived from stochastic group-wise rollouts to estimate prediction consistency as a proxy for model confidence. Together, these two components encourage the model to reinforce accurate predictions on clear cases while remaining cautious under ambiguity. Experiments across liver, breast, and thyroid datasets show that our framework improves lesion localization by 39.3\%, demonstrating that our model has learned the ability to actively look closer and diagnose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。