让大模型精准理解复杂指令并定位目标,突破现有分割推理瓶颈
StAR: Segment Anything Reasoner
- 通过多维度优化设计,激活基础模型的隐含推理能力
- 仅用5000样本训练即超越多个基线模型,性能显著提升
- 构建新基准ReasonSeg-X,支持深度推理任务评估
随着人工智能系统快速融入多样复杂的现实环境,基于隐式查询与图像进行整体推理以定位目标的能力愈发重要。然而,现有推理分割方法未能充分激发基础模型的视觉推理潜力。本文提出Segment Anything Reasoner(StAR),从参数调优策略、奖励函数、学习方式和答案格式等多个维度重构设计空间,显著优于近期基线方法。首次将并行测试时扩展引入分割任务,进一步突破性能边界。为拓展现有基准的覆盖范围与深度,我们构建了ReasonSeg-X数据集,精确定义推理类型,并包含需深层推理的样本。利用该数据集,采用滚动展开的选择性微调方法训练StAR,有效激活基础模型的潜在推理能力,并建立系统化、细粒度的评估基准。仅使用5000个训练样本,StAR在广泛基准上均取得显著提升,证明其能有效唤醒模型的沉睡推理能力。
原文摘要 · Abstract (English)
As AI systems are being integrated more rapidly into diverse and complex real-world environments, the ability to perform holistic reasoning over an implicit query and an image to localize a target is becoming increasingly important. However, recent reasoning segmentation methods fail to sufficiently elicit the visual reasoning capabilities of the base mode. In this work, we present Segment Anything Reasoner (StAR), a comprehensive framework that refines the design space from multiple perspectives-including parameter-tuning scheme, reward functions, learning strategies and answer format-and achieves substantial improvements over recent baselines. In addition, for the first time, we successfully introduce parallel test-time scaling to the segmentation task, pushing the performance boundary even further. To extend the scope and depth of reasoning covered by existing benchmark, we also construct the ReasonSeg-X, which compactly defines reasoning types and includes samples that require deeper reasoning. Leveraging this dataset, we train StAR with a rollout-expanded selective-tuning approach to activate the base model's latent reasoning capabilities, and establish a rigorous benchmark for systematic, fine-grained evaluation of advanced methods. With only 5k training samples, StAR achieves significant gains over its base counterparts across extensive benchmarks, demonstrating that our method effectively brings dormant reasoning competence to the surface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。