提出可解释的分割推理框架,让模型像人一样一步步思考。
GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine

- 将视觉理解、语义描述与大模型推理解耦,形成可追踪的逻辑链。
- 零样本下在多个分割基准上表现优异,构建了超大规模数据集GEAR-131K。
- 自动生成标注数据,轻量模型性能逼近人工标注上限。
推理分割需基于复杂隐含查询定位目标。现有端到端模型通常将感知与推理混为一锅黑箱,严重限制可解释性与可扩展性。为此,我们提出GEAR-Seg(面向推理分割的具身可解释智能体),一种显式解耦的智能体,通过将视觉像素转化为密集、属性丰富的文本,实现感知、描述与大语言模型(LLM)推理的分离。该框架将隐式推理转变为可追踪的显式逻辑链。作为零样本推理框架,其在多样化推理与细粒度指代分割基准上达到领先性能。此外,GEAR-Seg天然具备高可扩展的数据引擎功能。利用该引擎,我们构建了包含38,000+图像、656,000个问答-掩码对的巨型基准集GEAR-131K,引入多维度分类体系以支持复杂现实场景下的操作型推理。最后,蒸馏实验表明,仅由自动化流水线监督的轻量模型性能接近昂贵人工标注基线的上限。
原文摘要 · Abstract (English)
Reasoning segmentation requires localizing targets based on complex, implicit queries. Current end-to-end models typically entangle perception and deduction into an opaque black box, severely limiting interpretability and scalability. To address this, we propose GEAR-Seg (Grounded Explainable Agent for Reasoning Segmentation), an explicitly decoupled agent that shifts the paradigm by translating visual pixels into dense, attribute-rich text. By decoupling class-agnostic segmentation, semantic description, and Large Language Model (LLM) deduction, GEAR-Seg transforms implicit reasoning into an explicit, trackable logic chain. As a zero-shot inference framework, it achieves highly competitive performance across diverse reasoning and fine-grained referring segmentation benchmarks. Furthermore, GEAR-Seg inherently functions as a highly scalable data engine. Utilizing this engine, we construct GEAR-131K, a massive benchmark (over 38k images, 656k QA-mask pairs) introducing a multifaceted taxonomy tailored for complex real-world manipulation-oriented reasoning. Finally, distillation experiments demonstrate that lightweight models supervised exclusively by our automated pipeline closely match the upper-bound performance of costly human-annotated baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。