arXiv:2604.21914cs.RO2026-04中稿 · ICRA被引 3

让机器人在不同视角下都能稳定操作,无需校准相机。

VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis

论文配图:VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
图 1 · 摘自论文原文
  • 用几何模型与视频扩散结合生成多视角图像。
  • 跨视角泛化能力提升2.79倍以上,真实场景表现优异。
  • 适合需要多角度适应的机器人操作研究者使用。

近期端到端机器人操控模型因其泛化性和可扩展性受到关注,但固定摄像头训练时对视角变化鲁棒性差。本文提出VistaBot框架,通过融合前馈几何模型与视频扩散模型,实现无需测试时相机标定的视点鲁棒闭环操控。该方法包含四个关键组件:4D几何估计、视角合成潜在表示提取、潜在动作学习。VistaBot集成于动作分块(ACT)与基于扩散模型(π₀)的策略,在仿真和真实任务中均验证有效。我们引入新的跨视角泛化评分(View Generalization Score, VGS)以全面评估泛化性能。结果表明,VistaBot相较ACT和π₀分别提升VGS 2.79倍和2.63倍,同时实现高质量的新视角合成。贡献包括几何感知合成模型、潜在动作规划器、新基准评估指标及多环境验证。代码与模型将公开发布。

原文摘要 · Abstract (English)

Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that integrates feed-forward geometric models with video diffusion models to achieve view-robust closed-loop manipulation without requiring camera calibration at test time. Our approach consists of three key components: 4D geometry estimation, view synthesis latent extraction, and latent action learning. VistaBot is integrated into both action-chunking (ACT) and diffusion-based ($π_0$) policies and evaluated across simulation and real-world tasks. We further introduce the View Generalization Score (VGS) as a new metric for comprehensive evaluation of cross-view generalization. Results show that VistaBot improves VGS by 2.79$\times$ and 2.63$\times$ over ACT and $π_0$, respectively, while also achieving high-quality novel view synthesis. Our contributions include a geometry-aware synthesis model, a latent action planner, a new benchmark metric, and extensive validation across diverse environments. The code and models will be made publicly available.

机器人操控视角鲁棒视频生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。