arXiv:2505.22596cs.CV2025-05NeurIPS被引 31

用强化学习让多模态模型学会精细分割,仅需3000样本。

SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning

  • 用强化学习+细粒度奖励训练多模态模型,无需人工标注推理过程。
  • 仅用3000样本在多个基准上表现优异,超越传统依赖标注数据的方法。
  • 首次将SAM作为奖励函数用于细粒度分割训练,适合高效模型开发场景。

利用多模态大模型进行图像分割已成为研究热点。然而,现有方法通常严重依赖包含显式推理过程的人工标注数据,这类数据制作成本高、耗时长。最近进展表明,强化学习(RL)可在无需推理标注数据的情况下赋予大模型推理能力。本文提出SAM-R1,一种新框架,使多模态大模型能在图像理解任务中实现精细推理。我们的方法首次在多模态推理模型训练中引入细粒度分割设置,通过结合任务特定的细粒度奖励与定制优化目标,进一步提升模型推理与分割的一致性。我们还利用分隔一切模型(Segment Anything Model, SAM)作为强大且灵活的奖励提供者,引导学习过程。仅使用3000个训练样本,SAM-R1在多个基准测试中表现出色,验证了强化学习在赋予多模态模型面向分割的推理能力方面的有效性。

原文摘要 · Abstract (English)

Leveraging multimodal large models for image segmentation has become a prominent research direction. However, existing approaches typically rely heavily on manually annotated datasets that include explicit reasoning processes, which are costly and time-consuming to produce. Recent advances suggest that reinforcement learning (RL) can endow large models with reasoning capabilities without requiring such reasoning-annotated data. In this paper, we propose SAM-R1, a novel framework that enables multimodal large models to perform fine-grained reasoning in image understanding tasks. Our approach is the first to incorporate fine-grained segmentation settings during the training of multimodal reasoning models. By integrating task-specific, fine-grained rewards with a tailored optimization objective, we further enhance the model's reasoning and segmentation alignment. We also leverage the Segment Anything Model (SAM) as a strong and flexible reward provider to guide the learning process. With only 3k training samples, SAM-R1 achieves strong performance across multiple benchmarks, demonstrating the effectiveness of reinforcement learning in equipping multimodal models with segmentation-oriented reasoning capabilities.

多模态强化学习图像分割SAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。