arXiv:2605.19410cs.CV2026-05被引 1

无需训练的视觉代理,能自动构建复杂图像分割结果

Vision Harnessing Agent for Open Ad-hoc Segmentation

论文配图:Vision Harnessing Agent for Open Ad-hoc Segmentation
图 1 · 摘自论文原文
  • 用视觉工作记忆和规划能力逐步构造分割掩码
  • 在新基准上比现有方法高14-25%,多粒度引用分割提升5-9%
  • 适合需要动态理解新概念的视觉智能系统研究者

当概念已知时,分割变得简单,只需从文本中检索已学习的视觉定位。但对于开放的临时概念,定位可能不存在于单一已学习掩码中,需通过图像证据中的部分、关系、排除和集合来构建。我们提出首个面向开放临时分割的视觉引导代理VASA,其无需训练,结合视觉语言模型、分割基础模型与视觉化工作流程。VASA不只修改文本提示,而是使用持续的工作掩码进行推理、构建与验证,规划视觉操作,调用分割工具,检查结果,编辑掩码并处理错误。我们构建了PARS新基准,将PartImageNet中的部件级标签转化为通过长格式定义查询的开放临时概念。在PARS上,VASA优于开放词汇、基于推理及代理类基线,超越SAM3 Agent 14-25%;在RefCOCOm标准多粒度指代分割基准上,相比SAM3 Agent提升5-9%,较其他代理基线最高提升20%。这些结果验证了代理式视觉构造在开放临时分割中的有效性。本工作为超越封装基础模型作为工具的AI代理指明方向:赋予任务知识、视觉语言模型行为、视觉操作流程、工作记忆与容错工作流。

原文摘要 · Abstract (English)

Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-hoc concepts, where the grounding may not exist as one learned mask and must often be constructed from image evidence through parts, relations, exclusions, and collections. We propose a Vision-guided Ad-hoc Segmentation Agent (VASA), the first vision harnessing agent for open ad-hoc segmentation. VASA is training-free and couples a VLM agent, a segmentation foundation model, and a visually grounded workflow. Rather than revising text prompts alone, VASA uses a persistent working mask to reason, construct, and validate a solution. It plans visual operations, invokes segmentation tools, inspects results, edits the mask, and recovers from errors. We construct PARS, a new benchmark that turns part-level labels in PartImageNet into open ad-hoc concepts through long-form definition queries. On PARS, VASA outperforms open-vocabulary, reasoning-based, and agentic baselines, surpassing SAM3 Agent by 14-25%. On RefCOCOm, a standard multi-granularity referring segmentation benchmark, VASA improves over SAM3 Agent by 5-9% and over other agentic baselines by up to 20%. These results validate agentic visual construction for open ad-hoc segmentation. Our work points to a path for AI agents beyond wrapping foundation models as tools: Programming them with task knowledge, VLM behavior, visual routines, working memory, and failure-aware workflows.

图像分割视觉代理开放概念

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。