arXiv:2602.04880cs.RO2026-02被引 1

用环境状态解码能力评估视觉表征,能高效预测机器人控制性能。

Capturing Visual Environment Structure Correlates with Control Performance

  • 通过解码图像中的几何、物体结构和物理属性来评估视觉编码器
  • 解码准确率与下游策略性能强相关,优于现有指标
  • 适合研究通用机器人策略的表征学习与选型

视觉表征的选择对通用机器人策略的扩展至关重要。然而,通过策略回放直接评估代价高昂,即使在仿真中也是如此。现有代理指标仅关注视觉世界中狭窄方面的表征能力,如物体形状,限制了跨环境的泛化能力。本文从分析角度出发:通过测量预训练视觉编码器从图像中解码环境状态(包括几何、物体结构和物理属性)的能力来探测其性能。利用可访问真实状态的仿真环境,我们发现这种探测准确率在多种环境和学习设置下与下游策略性能高度相关,显著优于以往指标,并实现了高效的表征选择。更广泛地说,本研究揭示了支持可泛化操作的表征特性,表明学习编码环境潜在物理状态是一个有前景的控制目标。

原文摘要 · Abstract (English)

The choice of visual representation is key to scaling generalist robot policies. However, direct evaluation via policy rollouts is expensive, even in simulation. Existing proxy metrics focus on the representation's capacity to capture narrow aspects of the visual world, like object shape, limiting generalization across environments. In this paper, we take an analytical perspective: we probe pretrained visual encoders by measuring how well they support decoding of environment state -- including geometry, object structure, and physical attributes -- from images. Leveraging simulation environments with access to ground-truth state, we show that this probing accuracy strongly correlates with downstream policy performance across diverse environments and learning settings, significantly outperforming prior metrics and enabling efficient representation selection. More broadly, our study provides insight into the representational properties that support generalizable manipulation, suggesting that learning to encode the latent physical state of the environment is a promising objective for control.

机器人控制视觉表征状态解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。