arXiv:2601.23220cs.CVcs.AI2026-01中稿 · ICML被引 1

用几何感知强化学习,解决医疗多模态模型的视觉失真问题

Med-Scout: Curing MLLMs' Geometric Blindness in Medical Perception via Geometry-Aware RL Post-Training

  • 通过临床阅读逻辑设计三类自监督任务,从无标注医学图像中提取几何约束信号
  • 在新基准上使模型几何感知能力提升超40%,且在放射科问答中表现更优
  • 无需专家标注,适合需要高精度视觉理解的医疗AI研发人员

尽管多模态大模型在医疗诊断中表现出强大的语言能力,我们发现即使是顶尖模型也存在关键的感知缺陷:几何盲视。这种无法将输出锚定在客观几何约束上的问题,导致模型产生看似合理却事实错误的幻觉,根源在于训练范式过度追求语言流畅性而非几何保真度。本文提出Med-Scout框架,通过强化学习“治愈”这一盲视问题,利用未标注医学图像中隐含的几何逻辑作为监督信号。该方法基于临床系统的读片与推理模式,设计了三类代理任务:分层尺度定位、拓扑拼图重构和异常一致性检测。为严谨评估此缺陷,我们构建了专门用于测试几何感知能力的新基准Med-Scout-Bench。大量实验表明,Med-Scout显著缓解了几何盲视,在该基准上优于领先专有及开源模型超过40%;且增强的几何感知能力可泛化至更广泛的医疗理解任务,在放射科和综合医疗VQA任务中均取得更优结果。

原文摘要 · Abstract (English)

Despite recent Multimodal Large Language Models (MLLMs)' linguistic prowess in medical diagnosis, we find even state-of-the-art MLLMs suffer from a critical perceptual deficit: geometric blindness. This failure to ground outputs in objective geometric constraints leads to plausible yet factually incorrect hallucinations, rooted in training paradigms that prioritize linguistic fluency over geometric fidelity. This paper introduces Med-Scout, a novel framework that "cures" this blindness via Reinforcement Learning (RL) that leverages the intrinsic geometric logic latent within unlabeled medical images. Instead of relying on costly expert annotations, Med-Scout derives verifiable supervision signals through three strategic proxy tasks inspired by the systematic reading and reasoning patterns of clinicians: Hierarchical Scale Localization, Topological Jigsaw Reconstruction, and Anomaly Consistency Detection. To rigorously quantify this deficit, we present Med-Scout-Bench, a new benchmark specifically designed to evaluate geometric perception. Extensive evaluations show that Med-Scout significantly mitigates geometric blindness, outperforming leading proprietary and open-source MLLMs by over 40% on our benchmark. Furthermore, this enhanced geometric perception generalizes to broader medical understanding, achieving superior results on radiological and comprehensive medical VQA tasks.

医疗AI几何感知强化学习多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。