让医学影像大模型像医生一样动态复查不确定区域,生成更可靠的报告。
UR$^{2}$-MLLM: Uncertainty-aware Revisit Reasoning in Multimodal Large Language Models for Radiology Report Generation

- 引入不确定性感知机制,动态识别需复查的影像区域。
- 在MIMIC-CXR和IU-Xray上超越现有方法,报告准确率显著提升。
- 适合医学影像生成、临床辅助诊断等需要高可信度的场景。
放射科医生通过反复检查可疑区域来逐步完善诊断结论。当前用于放射学报告生成的多模态大模型虽转向‘图像思考’范式,但缺乏推理过程中动态回溯机制,未能模拟医生对不确定发现的反复审视。为此,我们提出不确定性感知的回溯推理框架UR$^{2}$-MLLM,实现报告生成中对不确定区域的动态重访。该模型首先在不确定性感知数据集上训练以获得感知能力;接着构建多模态推理轨迹数据集并引入检测-复制机制,指导何时何地进行回溯;最后通过视觉定位奖励,利用强化学习使回溯区域与解剖结构精准对齐。在MIMIC-CXR和IU-Xray数据集上的实验表明,UR$^{2}$-MLLM达到当前最佳性能,验证了不确定性感知视觉回溯推理对生成可靠且临床一致报告的价值。
原文摘要 · Abstract (English)
Radiologists generate diagnostic reports through iterative and selective revisiting of suspicious regions to refine their interpretations. Recent multimodal large language models (MLLMs) for radiology report generation (RRG) have shifted from text-only reasoning toward a ``Thinking-with-Images'' paradigm, incorporating visual evidence into the reasoning process. However, existing methods provide static visual evidence without a dynamic revisit mechanism during reasoning, neglecting how radiologists re-examine uncertain observations. To this end, we propose an Uncertainty-aware Revisit Reasoning MLLM (UR$^{2}$-MLLM) framework that dynamically revisits uncertain regions during reasoning for RRG. UR$^{2}$-MLLM is first equipped with uncertainty perception by training on an uncertainty-aware dataset. We then construct a multimodal reasoning trajectory dataset together with a detect-and-copy mechanism, which guides when and where to revisit. Finally, a visual grounding reward refines this behavior through reinforcement learning, aligning the revisited regions with corresponding anatomical structures. Experiments on MIMIC-CXR and IU-Xray show that UR$^{2}$-MLLM achieves state-of-the-art performance, highlighting the value of uncertainty-aware visual revisit reasoning for reliable and clinically aligned report generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。