用双摄像头预测多视角未来,让机器人看清被遮挡的物体
UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling

- 从两视角图像生成多视角未来画面,建模场景演化
- 遮挡任务成功率从40%提升至73.3%,真实机器人性能提高33.4点
- 无需额外相机或3D重建,适合复杂遮挡环境下的机器人控制
遮挡任务仍是机器人操作的瓶颈。现有方法要么需额外物理摄像头(要求训练推理视角一致),要么依赖高计算成本的显式3D重建。两者均仅使用标准的主视角和腕部视角观测,无法捕捉遮挡信息与未来场景演变。为此,我们提出UniviewVLA——一种统一的多视角视觉-语言-动作模型,结合世界建模,仅凭标准双摄像头观测即可推断多视角未来场景演化。通过利用世界模型生成的多视角未来视图,UniviewVLA揭示被遮挡线索并建模未来演化,提升动作预测能力,同时避免额外硬件或显式重建。此外,为加速推理,UniviewVLA引入运动感知令牌压缩技术,将每幅生成视图的令牌数从625压缩至16,单视图延迟由6–7秒降至0.2–0.3秒;提出免训练的动作熵视图选择机制,动态识别不同推理阶段最动作相关的视图。大量实验表明,UniviewVLA在LIBERO上达95.8%,在CALVIN ABCD to D上达4.60,均为标准无遮挡基准。在定制遮挡任务中,成功率达73.3%(原40.0%),真实机器人平均成功率提升33.4点,展现更强遮挡适应能力,且不牺牲标准基准表现。
原文摘要 · Abstract (English)
Occluded tasks remain a bottleneck in robot manipulation. Existing solutions either deploy additional physical cameras requiring training-inference camera parity, or rely on explicit 3D reconstruction with high computational cost. Moreover, both approaches rely on standard agent-view and wrist-view observations, while failing to capture occlusion information and future scene evolution. To this end, we propose UniviewVLA, a unified multiview Vision-Language-Action model with world modeling, which infers multiview scene evolution for action prediction from only standard two-camera observations. We demonstrate that by leveraging generated multiview future views from the world model, UniviewVLA reveals occluded cues and models future scene evolution, improving action prediction and removing the need for extra hardware or explicit reconstruction. Besides, to accelerate inference while preserving prediction accuracy, UniviewVLA develops Motion-Informative Token Compression, which compresses each generated view from 625 to 16 tokens and reduces per-view latency from 6-7s to 0.2-0.3s. UniviewVLA also proposes training-free Action-Entropy View Selection, which dynamically identifies the most action-informative view at different inference stages. Extensive experiments show that UniviewVLA achieves 95.8% on LIBERO and 4.60 on CALVIN ABCD to D, both standard occlusion-free benchmarks. On customized occlusion-focused tasks, it improves success rate from 40.0% to 73.3%, and average real-robot success rate by 33.4 points, demonstrating stronger occlusion-focused performance without sacrificing standard occlusion-free benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。