arXiv:2503.06520cs.CVcs.MM2025-03被引 222

无需标注数据,通过思维链推理实现零样本图像分割

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

  • 分离推理与分割模块,用思维链生成定位提示
  • 仅用强化学习训练,零样本下达57.5分(超越前代18%)
  • 适合需要可解释分割的视觉任务研究者

传统推理分割方法依赖类别标签和简单描述进行监督微调,限制了其跨领域泛化能力且缺乏显式推理过程。为此,我们提出Seg-Zero,一种通过认知强化实现显著泛化性并生成显式思维链推理的新框架。该框架采用解耦架构,包含推理模型与分割模型:推理模型解析用户意图、生成显式推理链并输出位置提示,由分割模型据此生成精确像素级掩码。设计了融合格式与准确率奖励的复杂奖励机制,有效引导优化方向。仅通过强化学习(GRPO)训练,无需显式推理数据,Seg-Zero展现出强大的零样本泛化能力与涌现的测试时推理能力。实验表明,Seg-Zero-7B在ReasonSeg基准上达到57.5的零样本性能,超越先前LISA-7B 18%,凸显其跨域泛化能力与显式推理优势。

原文摘要 · Abstract (English)

Traditional methods for reasoning segmentation rely on supervised fine-tuning with categorical labels and simple descriptions, limiting its out-of-domain generalization and lacking explicit reasoning processes. To address these limitations, we propose Seg-Zero, a novel framework that demonstrates remarkable generalizability and derives explicit chain-of-thought reasoning through cognitive reinforcement. Seg-Zero introduces a decoupled architecture consisting of a reasoning model and a segmentation model. The reasoning model interprets user intentions, generates explicit reasoning chains, and produces positional prompts, which are subsequently used by the segmentation model to generate precious pixel-level masks. We design a sophisticated reward mechanism that integrates both format and accuracy rewards to effectively guide optimization directions. Trained exclusively via reinforcement learning with GRPO and without explicit reasoning data, Seg-Zero achieves robust zero-shot generalization and exhibits emergent test-time reasoning capabilities. Experiments show that Seg-Zero-7B achieves a zero-shot performance of 57.5 on the ReasonSeg benchmark, surpassing the prior LISA-7B by 18\%. This significant improvement highlights Seg-Zero's ability to generalize across domains while presenting an explicit reasoning process.

图像分割思维链强化学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。