arXiv:2603.05868cs.RO2026-03中稿 · IROS 2026被引 5

无需微调,实时适配相机视角,提升机器人视觉语言动作模型鲁棒性。

AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models

  • 通过实时生成目标视角图像,虚拟调整测试时相机视图。
  • 在LIBERO基准上超越基于数据增强或3D特征的方法,性能更稳定。
  • 适用于任意RGB策略,支持手持相机等复杂场景,部署零成本。

尽管视觉-语言-动作模型(VLAs)在机器人操作中取得显著进展,但其部署需针对特定环境微调,且对未结构化环境中频繁变化的相机视角极为敏感。本文提出一种零样本相机适配框架,无需额外演示数据、策略微调或架构修改。核心思想是实时将测试时的相机观测虚拟调整至训练时的配置。为此,我们采用近期前馈式新视角合成模型,可高质量输出目标视角图像,同时处理内外参变化。该即插即用方法保持预训练模型能力,适用于任何基于RGB的策略。在LIBERO基准上的大量实验表明,该方法持续优于使用数据增强微调或引入额外3D感知特征的基线方法。进一步验证显示,该方法在真实机器人操作场景中显著提升视角鲁棒性,涵盖相机外参、内参变化及自由移动手持相机等复杂设置。

原文摘要 · Abstract (English)

Despite remarkable progress in Vision-Language-Action models (VLAs) for robot manipulation, these large pre-trained models require fine-tuning to be deployed in specific environments. These fine-tuned models are highly sensitive to camera viewpoint changes that frequently occur in unstructured environments. In this paper, we propose a zero-shot camera adaptation framework without additional demonstration data, policy fine-tuning, or architectural modification. Our key idea is to virtually adjust test-time camera observations to match the training camera configuration in real-time. For that, we use a recent feed-forward novel view synthesis model which outputs high-quality target view images, handling both extrinsic and intrinsic parameters. This plug-and-play approach preserves the pre-trained capabilities of VLAs and applies to any RGB-based policy. Through extensive experiments on the LIBERO benchmark, our method consistently outperforms baselines that use data augmentation for policy fine-tuning or additional 3D-aware features for visual input. We further validate that our approach constantly enhances viewpoint robustness in real-world robotic manipulation scenarios, including settings with varying camera extrinsics, intrinsics, and freely moving handheld cameras. Project Page: https://heo0224.github.io/AnyCamVLA/

视觉语言动作相机适配零样本机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。