让机器人在不同视角下稳定执行动作,靠的是捕捉物理运动规律的隐式动作表示。
Learning to Act Robustly with View-Invariant Latent Actions
- 用轨迹中的动作模式建模隐式动作,融合物理动态而非仅依赖外观
- 在仿真和真实场景中均实现未见视角下的有效泛化
- 适合需要强鲁棒性的机器人控制任务预训练
基于视觉的机器人策略常因视角微小变化而失效,亟需视图不变的视觉表征。现实场景中视角变化不可避免,严重影响策略性能。现有方法多在场景层面利用多视角观测学习不变性,但依赖视觉外观,忽视了对鲁棒泛化至关重要的物理动态。本文提出视图不变隐式动作(VILA),通过建模跨轨迹的隐式动作以捕捉转移模式,从而基于物理动态学习视图不变表征。VILA采用基于真实动作序列的动作引导目标,在不同视角间对齐隐式动作。仿真与真实世界实验表明,基于VILA的策略能有效泛化至未见视角,并良好迁移到新任务,验证其作为强预训练框架在提升鲁棒性和下游学习性能方面的有效性。
原文摘要 · Abstract (English)
Vision-based robotic policies often struggle with even minor viewpoint changes, underscoring the need for view-invariant visual representations. This challenge becomes more pronounced in real-world settings, where viewpoint variability is unavoidable and can significantly disrupt policy performance. Existing methods typically learn invariance from multi-view observations at the scene level, but such approaches rely on visual appearance and fail to incorporate the physical dynamics essential for robust generalization. We propose View-Invariant Latent Action (VILA), which models a latent action capturing transition patterns across trajectories to learn view-invariant representations grounded in physical dynamics. VILA aligns these latent actions across viewpoints using an action-guided objective based on ground-truth action sequences. Experiments in both simulation and the real world show that VILA-based policies generalize effectively to unseen viewpoints and transfer well to new tasks, establishing VILA as a strong pretraining framework that improves robustness and downstream learning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。