arXiv:2606.31599cs.CVcs.AI2026-06

通过双流强化学习,让医疗多模态模型只关注关键图像区域,提速提效。

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning

论文配图:Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning
图 1 · 摘自论文原文
  • 用双分支强化学习,一边定位关键图像区域,一边做稀疏推理。
  • 在7个医疗数据集上将图像标记数减少77%,性能提升超100%。
  • 适合追求高效医疗视觉推理的科研与临床应用者。

融合强化学习的视觉语言模型在多模态推理中取得显著进展,但在医疗图像场景中仍表现不足,因医学图像通常仅含极少的视觉证据支持临床决策。我们发现,剔除非关键区域的视觉标记可显著提升医疗推理效果。然而,尚无统一的强化学习框架实现主动视觉标记剪枝(VTP)与多模态推理协同优化。为此,我们提出双流强化学习框架ViToS,通过一个共享策略模型的双任务分支,分别聚焦于视觉定位与剪枝后的稀疏推理。为解决策略耦合问题,引入跨反馈序列优化机制,避免梯度冲突并促进收敛。在七个医疗基准上评估,该方法将视觉标记长度压缩至原长的77%,在Lingshu-7B上实现108.27%相对性能提升,在HuatuoGPT-Vision-7B上达104.16%相对提升。整体表现更优且推理速度加快,建立了高效的医疗多模态推理新范式。

原文摘要 · Abstract (English)

Vision-language models (VLMs) combining reinforcement learning (RL) ignite remarkable progress in multimodal reasoning, yet still struggle with medical images, which typically exhibit extremely sparse visual evidence to inform clinical decision-making. We recognize that pruning visual tokens outside the grounding region greatly enhances medical reasoning. However, a united RL framework for active visual token pruning (VTP) and medical multimodal reasoning remains unestablished. Here, we propose a dual-stream RL framework, ViToS, to fulfill token pruning and question answering. ViToS trains one policy model with two task branches, where one focuses on grounding while the other conducts token-sparse reasoning after VTP. Furthermore, we solve the coupled policy learning problem by introducing the cross-feedback sequential optimization, avoiding gradient conflict and facilitating convergence of the shared policy model. Evaluated on seven medical benchmarks, our method reduces visual tokens to 77% of the original sequence length while achieving a 108.27% relative performance on Lingshu-7B and 104.16% relative performance on HuatuoGPT-Vision-7B. Overall, ViToS delivers superior performance and inference speedup, establishing an efficient paradigm for medical multimodal reasoning.

医疗多模态强化学习视觉推理高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。