arXiv:2608.06965cs.RO2026-08被引 1

让机器人在不同摄像头视角下仍能准确执行任务,提升视觉-语言-动作策略的鲁棒性。

Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies

论文配图:Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies
图 1 · 摘自论文原文
  • 通过构建动作等价的视图对,强制模型关注场景本身而非摄像头位置。
  • 在真实机器人上,跨摄像头任务成功率从53.3%提升至74.4%。
  • 无需深度或相机参数,适合实际部署中的多视角应用。

从固定摄像头场景微调的视觉-语言-动作(VLA)策略在摄像头移动时会失效,即使任务、物体、语言和机器人状态不变。本文仅使用场景RGB图像、语言和本体感知信息,不依赖摄像头标签、外参、深度或点云输入,研究场景视角鲁棒性。为防止手腕视觉流成为未扰动的捷径,全程屏蔽腕部流。针对基于光流的VLA,提出正则化动作光流速度场——直接积分生成连续动作块的核心量。通过重置原始LIBERO演示至同一MuJoCo状态,并渲染名义与扰动摄像头视角,构建动作等价视图对。双视图均通过光流匹配监督,同时引入跨视图损失,使相同采样坐标处预测的动作光流速度一致。在LIBERO-Plus摄像头扰动赛道上,本方法达87.2±0.4%(每种子3次训练共4,797次回放),较仅用光流匹配提升7.4个百分点(79.8±0.8%,同数据集3种子),较朴素混合摄像头微调提升12.5个百分点;保持名义摄像头性能95.0±0.8%(同数据集光流匹配:95.0±4.3%)。随机配对对照组下降至25.8%,证明增益依赖于动作等价配对。在真实机器人上,3个桌面任务各10次回放,跨摄像头成功率从53.3%提升至74.4%,使用相同的单场景RGB推理接口。

原文摘要 · Abstract (English)

Vision-language-action (VLA) policies fine-tuned from a fixed scene camera can fail when the camera is moved, even when the task, objects, language, and robot state are unchanged. We study scene-camera viewpoint robustness using only a scene RGB image, language, and proprioception, without camera labels, extrinsics, depth, or point-cloud inputs. The wrist stream is masked throughout to prevent an unperturbed visual shortcut from confounding attribution to scene-camera variation. For flow-based VLAs, we propose to regularize the action-flow velocity field, the quantity directly integrated to generate continuous action chunks. We construct action-equivalent view pairs by resetting original LIBERO demonstrations to the same MuJoCo state and rendering nominal and perturbed scene-camera views. Both views are supervised by flow matching, while a cross-view loss encourages their predicted action-flow velocities to agree at the same sampled flow coordinates. On the LIBERO-Plus camera-perturbation track, our method reaches 87.2$\pm$0.4% (4,797 rollouts per seed across 3 training seeds), +7.4pp over flow-matching-only training on the same paired data (79.8$\pm$0.8%, also 3 seeds) and +12.5pp over naive mixed-camera SFT, while maintaining nominal-camera ID performance (95.0$\pm$0.8%; same-data FM-only: 95.0$\pm$4.3%). A shuffled-pair control collapses to 25.8%, showing that the gain depends on action-equivalent pairing. On a real robot, we evaluate three tabletop tasks with 10 rollouts per task and camera placement; held-out-camera success improves from 53.3% to 74.4% under the same single-scene-RGB inference interface.

视觉-语言-动作多视角鲁棒机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。