arXiv:2603.13904cs.CVcs.AI2026-03中稿 · CVPR

让视觉状态同时记住物体是什么、在哪,提升机器人对动态环境的理解能力

Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition

  • 用全局-局部重建目标,让单个令牌编码场景的语义与位置信息
  • 在多个机器人任务上达到当前最优,能精准捕捉物体移动与交互变化
  • 适合需要理解像素级场景变化的机器人感知与决策系统

对于在动态环境中运行的机器人代理,从流式视频观测中学习视觉状态表示是实现序列决策的关键。近期自监督学习方法在跨视觉任务间表现出强泛化能力,但未明确说明良好视觉状态应包含什么。我们认为,有效的视觉状态必须通过联合编码场景元素的语义身份及其空间位置,实现对观测间微小动态的可靠检测。为此,我们提出 CroBo,一种基于全局到局部重建目标的视觉状态表示学习框架。给定一个压缩为紧凑瓶颈令牌的参考观测,CroBo 学习利用全局瓶颈令牌作为上下文,从稀疏可见线索中重建局部目标区域内的大量遮蔽补丁。该学习目标促使瓶颈令牌编码场景级语义实体的细粒度表示,包括其身份、空间位置及配置。结果,所学视觉状态揭示了场景元素随时间的运动与交互,支持序列决策。我们在多种基于视觉的机器人策略学习基准上评估 CroBo,表现达当前最优。重建分析与感知直线性实验进一步表明,所学表示保持像素级场景组合,并编码跨观测的‘什么在动、在哪动’信息。

原文摘要 · Abstract (English)

For robotic agents operating in dynamic environments, learning visual state representations from streaming video observations is essential for sequential decision making. Recent self-supervised learning methods have shown strong transferability across vision tasks, but they do not explicitly address what a good visual state should encode. We argue that effective visual states must capture what-is-where by jointly encoding the semantic identities of scene elements and their spatial locations, enabling reliable detection of subtle dynamics across observations. To this end, we propose CroBo, a visual state representation learning framework based on a global-to-local reconstruction objective. Given a reference observation compressed into a compact bottleneck token, CroBo learns to reconstruct heavily masked patches in a local target crop from sparse visible cues, using the global bottleneck token as context. This learning objective encourages the bottleneck token to encode a fine-grained representation of scene-wide semantic entities, including their identities, spatial locations, and configurations. As a result, the learned visual states reveal how scene elements move and interact over time, supporting sequential decision making. We evaluate CroBo on diverse vision-based robot policy learning benchmarks, where it achieves state-of-the-art performance. Reconstruction analyses and perceptual straightness experiments further show that the learned representations preserve pixel-level scene composition and encode what-moves-where across observations. Project page available at: https://seokminlee-chris.github.io/CroBo-ProjectPage.

视觉状态机器人感知自监督学习空间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。