用视觉空间结构化推理提升感知策略学习,避免语言模糊性
Artemis: Structured Visual Reasoning for Perception Policy Learning
- 中间步骤用(标签, 边界框)对表示,实现可验证的视觉状态追踪
- 在自然图像上训练后,可泛化到计数与几何感知任务
- 无需特定任务设计,统一架构支持多种感知任务
当前基于强化学习的视觉感知策略框架通常采用自然语言表达的中间推理链。实证发现,纯语言形式的推理常导致感知任务性能下降。我们认为问题不在推理本身,而在于推理形式:现有方法在无结构的语言空间中进行语义推理,但视觉感知需在空间和以对象为中心的空间中推理。为此,我们提出Artemis,一种执行结构化视觉推理的感知策略学习方法,其中每个中间步骤以(标签, 边界框)对形式表示,捕捉可验证的视觉状态。该设计支持中间状态的显式追踪、提议质量的直接监督,并避免语言推理引入的歧义。基于可验证且空间锚定的推理链,Artemis为多种感知任务提供统一架构,无需依赖以往模型的任务特异性设计。在自然图像域的定位与检测样本上训练后,Artemis可泛化至计数与几何感知任务。其核心是空间锚定、以对象为中心的链式规则,为可扩展、通用的感知策略提供了原则性基础。
原文摘要 · Abstract (English)
Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural language. Empirical observations indicate that such purely linguistic intermediate reasoning often reduces performance on perception tasks. We argue that the core issue lies not in reasoning per se but in the form of reasoning: while these chains perform semantic reasoning in an unstructured linguistic space, \textbf{visual perception requires reasoning in a spatial and object-centric space}. In response, we introduce \textbf{Artemis}, a perception-policy learning method that performs structured visual reasoning, where each intermediate step is represented as a (label, bounding-box) pair capturing a verifiable visual state. This design enables explicit tracking of intermediate states, direct supervision for proposal quality, and avoids ambiguity introduced by language-based reasoning. Building upon verifiable and spatially grounded reasoning chains, Artemis provides a unified architecture for diverse perceptual tasks, without requiring the task-specific designs relied upon by prior perceptual policy models. Trained using grounding and detection sampeles in natural image domains, Artemis generalizes to counting and geometric perception tasks. At its core, a spatially grounded, object-centric chain rule provides a principled foundation for scalable and general perceptual policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。