arXiv:2509.18183cs.CVcs.AI2025-09被引 1

让机器人更灵活地应对不同视角的指令,提升复杂环境下的操作成功率。

VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation

  • 用轻量级模块融合多视角图像特征,实现视角自适应。
  • 在多个数据集上平均提升8%-30%的任务成功率。
  • 适合需要多视角感知的现实世界机器人任务。

视觉-语言-动作(VLA)模型能根据环境视觉观测执行文本指令,其能力源于大量标准示范训练。然而,第三视角全局相机与腕部局部相机拍摄的视觉观测在视角和数量上差异显著,导致视觉特征不一致,限制了VLA模型的泛化性。为此,我们提出轻量级模块VLA-LPAF,仅使用2D图像即可增强VLA模型的视角自适应能力。该模块通过单视图图像微调,在潜在空间融合多视角观测,有效缓解视角不一致带来的问题。我们将VLA-LPAF集成至RoboFlamingo模型,构建RoboFlamingo-LPAF。实验表明,该模型在CALVIN、LIBERO及定制仿真基准上平均分别提升8%、15%和30%的任务成功率。真实场景任务也验证了其良好的视角自适应能力。

原文摘要 · Abstract (English)

The Visual-Language-Action (VLA) models can follow text instructions according to visual observations of the surrounding environment. This ability to map multimodal inputs to actions is derived from the training of the VLA model on extensive standard demonstrations. These visual observations captured by third-personal global and in-wrist local cameras are inevitably varied in number and perspective across different environments, resulting in significant differences in the visual features. This perspective heterogeneity constrains the generality of VLA models. In light of this, we first propose the lightweight module VLA-LPAF to foster the perspective adaptivity of VLA models using only 2D data. VLA-LPAF is finetuned using images from a single view and fuses other multiview observations in the latent space, which effectively and efficiently bridge the gap caused by perspective inconsistency. We instantiate our VLA-LPAF framework with the VLA model RoboFlamingo to construct RoboFlamingo-LPAF. Experiments show that RoboFlamingo-LPAF averagely achieves around 8% task success rate improvement on CALVIN, 15% on LIBERO, and 30% on a customized simulation benchmark. We also demonstrate the developed viewadaptive characteristics of the proposed RoboFlamingo-LPAF through real-world tasks.

机器人操作多视角融合视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。