arXiv:2608.01826cs.RO2026-08

多视角统一相机场让机器人用单色摄像机实现精准动作感知

Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies

论文配图:Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies
图 1 · 摘自论文原文
  • 构建跨视角共享的动作朝向潜在场,统一多相机观测
  • 在LIBERO上达98.9%成功率,提升22.4分,六类任务平均提升23.3分
  • 仅需RGB输入,部署无额外计算开销,适合真实机器人应用

视觉-语言-动作(VLA)模型在机器人操作中展现强大泛化能力,但复杂接触任务常需多相机协同捕捉末端执行器、物体与目标,克服遮挡问题。现有方法多简单拼接视图标记,导致动作表征缺乏度量深度且跨视角不一致。本文提出多视角统一相机场(MVUCF),一种仅用于训练的框架,通过坐标查询深度目标恢复可度量深度,借助预处理感知对应目标对齐不同相机观测同一物理点的标记。二者直接塑造动作模块所用隐藏状态。几何注入后移除深度、相机标定及辅助头,部署时仅使用原始RGB图像,无额外推理算力开销。持留测试验证更强深度恢复与跨视图匹配能力。在匹配的GR00T-N1.6设置下,MVUCF在LIBERO上达到98.9%,使LIBERO-Plus提升22.4分,并在六个涵盖触碰、移动放置与接触交互三类动作的RoboTwin任务中平均成功率提升23.3分。真实人形机器人实验进一步证明其在仅用RGB输入下的实际有效性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation, yet complex contact-rich tasks often benefit from multi-camera observations that jointly capture the end effector, objects, and targets under occlusion. Existing multi-camera VLAs usually concatenate view tokens, leaving action representations weak in metric depth and inconsistent across cameras. We introduce Multi-View Unified Camera Fields (MVUCF), a training-only framework that forms a shared action-facing latent field across views. A coordinate-query depth objective makes metric depth recoverable, while a preprocessing-aware correspondence objective aligns tokens observing the same physical point from different cameras. Both directly shape the hidden states consumed by the action module. After geometry injection, depth, camera calibration, and auxiliary heads are removed, so deployment uses the original RGB-only graph with no extra inference FLOPs. Held-out probes confirm stronger depth recovery and cross-view matching. Under matched GR00T-N1.6 settings, MVUCF reaches 98.9% on LIBERO, improves LIBERO-Plus by 22.4 points, and raises success by 23.3 points across six RoboTwin tasks spanning three action families: touch, move-and-place, and contact interaction. Real-world humanoid experiments further provide evidence of its practical effectiveness under RGB-only deployment.

多视角感知机器人控制视觉-语言-动作端到端学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。