arXiv:2607.02322cs.ROcs.CV2026-07

用移动相机+固定视角混合数据,提升机器人视觉模型的空间泛化能力

The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection

论文配图:The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection
图 1 · 摘自论文原文
  • 双机械臂协作:一臂操作,一臂移动拍图,打破视角依赖
  • 混合策略比纯静态或多视角更优,显著降低错误关联
  • 对多种模型都有效,适合需要强空间推理的机器人任务

视觉-语言-动作(VLA)模型在泛化机器人操作中表现优异,但其空间泛化能力仍脆弱。我们指出,单纯增加视角数量不足以解决问题,模型容易陷入捷径学习,依赖物体间固定相对姿态或相机与机器人基座的固定关系等虚假相关性,而非真正学习空间关系。为此,我们提出一种以数据为中心的解决方案:采用双臂结构,一臂执行操作,另一臂作为可移动环境相机。系统评估了三种数据分布模式:固定视角、多固定视角和移动视角。结果表明,结合连续相机运动与多样化静态视角的混合策略效果最佳,能显著减少虚假相关性同时保持训练稳定。实验显示,该策略有效缓解虚假相关性,使VLA模型可在未见过的相机位姿和物体配置下实现泛化,而仅增加静态视角的方法则失败。关键发现是,对捷径学习的敏感性和空间泛化困难是多种架构的共性特征。因此,所有测试模型(ACT、Diffusion及包含Pi0和Gr00t的VLA模型)均显著受益于我们的混合数据策略。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown remarkable promise in generalized robotic manipulation. However, their spatial generalization remains fragile. We argue that simply increasing the number of viewpoints is insufficient. Models often fall into the trap of Shortcut Learning, latching onto spurious correlations (e.g., fixed relative poses between objects or between the camera and robot base) rather than learning true spatial relationships. In this work, we propose a data-centric solution to enhance VLA spatial generalization. We utilize a dual-arm setup where one arm performs manipulation while the other serves as a mobile environmental camera. We systematically evaluate three data distribution patterns: Fixed, Multi-Fixed, and Moving Views. Our findings reveal that a hybrid strategy, combining continuous camera motion with diverse static viewpoints, yields the best performance by substantially reducing spurious correlations while maintaining training stability. Our experiments demonstrate that this strategy mitigates spurious correlations, enabling VLAs to generalize to unseen camera poses and object configurations where simply adding more static viewpoints fails. Crucially, we reveal that the susceptibility to shortcut learning and the struggle with spatial generalization are universal characteristics shared across diverse architectures. Consequently, all evaluated models (ACT, Diffusion, and VLA models including Pi0 and Gr00t) benefit significantly from our mixed data strategy.

机器人操作空间泛化数据收集视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。