arXiv:2509.18644cs.ROcs.AI2025-09被引 13

不用本体感觉信息,纯视觉也能让机器人更灵活地完成操作任务。

Do You Need Proprioceptive States in Visuomotor Policies?

  • 摒弃本体感觉输入,仅靠视觉观察预测动作
  • 垂直和水平空间泛化成功率分别提升至85%和64%
  • 适合需要跨机器人部署和数据效率高的真实场景

基于模仿学习的视觉-运动策略广泛用于机器人操作,通常结合视觉观测与本体感觉状态以实现精确控制。然而,本研究发现这种做法使策略过度依赖本体感觉输入,导致对训练轨迹过拟合,空间泛化能力差。为此,提出无状态策略(State-free Policy),完全移除本体感觉输入,仅根据视觉观测预测动作。该策略在相对末端执行器动作空间中构建,并依赖双广角腕部摄像头提供完整的任务相关视觉信息。实验证明,无状态策略在真实任务中表现显著更优:在抓取放置、复杂衬衫折叠及全身协同操作等任务中,跨不同机器人形态,在高度泛化上成功率从0%提升至85%,水平泛化从6%提升至64%。此外,该策略在数据效率和跨形态适应性方面也具优势,更具实际部署价值。

原文摘要 · Abstract (English)

Imitation-learning-based visuomotor policies have been widely used in robot manipulation, where both visual observations and proprioceptive states are typically adopted together for precise control. However, in this study, we find that this common practice makes the policy overly reliant on the proprioceptive state input, which causes overfitting to the training trajectories and results in poor spatial generalization. On the contrary, we propose the State-free Policy, removing the proprioceptive state input and predicting actions only conditioned on visual observations. The State-free Policy is built in the relative end-effector action space, and should ensure the full task-relevant visual observations, here provided by dual wide-angle wrist cameras. Empirical results demonstrate that the State-free policy achieves significantly stronger spatial generalization than the state-based policy: in real-world tasks such as pick-and-place, challenging shirt-folding, and complex whole-body manipulation, spanning multiple robot embodiments, the average success rate improves from 0% to 85% in height generalization and from 6% to 64% in horizontal generalization. Furthermore, they also show advantages in data efficiency and cross-embodiment adaptation, enhancing their practicality for real-world deployment. Discover more by visiting: https://statefreepolicy.github.io.

视觉控制机器人操作泛化能力无状态策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。