arXiv:2602.03320cs.CVcs.AI2026-02被引 6

让医学图像分割像人一样多轮思考,更准更省力。

MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement Learning

  • 用多轮交互+动态反馈训练智能体,模仿医生分步判断。
  • 在21个数据集上表现超越现有方法,减少冗余操作。
  • 适合医疗影像开发、智能诊断系统研究者使用。

医学图像分割正从专用模型转向通用框架。近期研究利用多模态大语言模型作为自主智能体,结合可验证奖励的强化学习(RLVR)调度如SAM等专业工具。但这些方法通常依赖单轮固定交互,缺乏过程监督,难以发挥交互工具的动态潜力,导致重复操作。为此,我们提出MedSAM-Agent,将交互分割重构为多步自主决策过程。首先设计混合提示策略生成专家级轨迹,使模型内化人类决策启发与自适应优化策略;其次构建两阶段训练流程,融合多轮端到端结果验证与临床真实感过程奖励,提升交互简洁性与决策效率。在6种医学模态、21个数据集上的实验表明,MedSAM-Agent达到当前最优性能,有效融合自主医疗推理与稳健迭代优化。代码已开源。

原文摘要 · Abstract (English)

Medical image segmentation is evolving from task-specific models toward generalizable frameworks. Recent research leverages Multi-modal Large Language Models (MLLMs) as autonomous agents, employing reinforcement learning with verifiable reward (RLVR) to orchestrate specialized tools like the Segment Anything Model (SAM). However, these approaches often rely on single-turn, rigid interaction strategies and lack process-level supervision during training, which hinders their ability to fully exploit the dynamic potential of interactive tools and leads to redundant actions. To bridge this gap, we propose MedSAM-Agent, a framework that reformulates interactive segmentation as a multi-step autonomous decision-making process. First, we introduce a hybrid prompting strategy for expert-curated trajectory generation, enabling the model to internalize human-like decision heuristics and adaptive refinement strategies. Furthermore, we develop a two-stage training pipeline that integrates multi-turn, end-to-end outcome verification with a clinical-fidelity process reward design to promote interaction parsimony and decision efficiency. Extensive experiments across 6 medical modalities and 21 datasets demonstrate that MedSAM-Agent achieves state-of-the-art performance, effectively unifying autonomous medical reasoning with robust, iterative optimization. Code is available \href{https://github.com/CUHK-AIM-Group/MedSAM-Agent}{here}.

医学图像智能体强化学习交互分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。