arXiv:2603.05911cs.CVcs.AI2026-03被引 1

用强化学习让模型像医生一样推理复杂病灶分割

CORE-Seg: Reasoning-Driven Segmentation for Complex Lesions via Reinforcement Learning

  • 引入思维链提示适配器,让模型边推理边分割
  • 在14K病灶数据集上达37.06%平均骰子系数,领先第二名14.89%
  • 适合需要可解释性医疗视觉分析的研究者

医学图像分割正从传统视觉模式匹配转向认知推理分析。尽管多模态大语言模型(MLLMs)在融合语言与视觉知识方面展现潜力,但仍存在显著差距:现有通用MLLM具备广泛常识但缺乏复杂病灶所需的专门视觉推理能力,而传统分割模型虽擅长像素级分割却缺乏逻辑可解释性。本文提出首个针对推理驱动复杂病灶分割的链式思维(CoT)基准ComLesion-14K。为此,我们设计了端到端的CORE-Seg框架,通过语义引导提示适配器将推理与分割结合,并采用从SFT到GRPO的渐进训练策略,配备自适应双粒度奖励机制以缓解奖励稀疏问题。该方法在测试中取得37.06%的平均骰子系数(较第二好基线高出14.89%),失败率降至18.42%。

原文摘要 · Abstract (English)

Medical image segmentation is undergoing a paradigm shift from conventional visual pattern matching to cognitive reasoning analysis. Although Multimodal Large Language Models (MLLMs) have shown promise in integrating linguistic and visual knowledge, significant gaps remain: existing general MLLMs possess broad common sense but lack the specialized visual reasoning required for complex lesions, whereas traditional segmentation models excel at pixel-level segmentation but lack logical interpretability. In this paper, we introduce ComLesion-14K, the first diverse Chain-of-Thought (CoT) benchmark for reasoning-driven complex lesion segmentation. To accomplish this task, we propose CORE-Seg, an end-to-end framework integrating reasoning with segmentation through a Semantic-Guided Prompt Adapter. We design a progressive training strategy from SFT to GRPO, equipped with an adaptive dual-granularity reward mechanism to mitigate reward sparsity. Our Method achieves state-of-the-art results with a mean Dice of 37.06\% (14.89\% higher than the second-best baseline), while reducing the failure rate to 18.42\%. Project Page: https://xyxl024.github.io/CORE-Seg.github.io/

医学图像推理分割强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。