arXiv:2606.18375cs.RO2026-06被引 3

让机器人世界模型实现多视角3D一致性,提升操作精度。

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

论文配图:PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation
图 1 · 摘自论文原文
  • 引入显式跨视角通信与3D几何先验,解决多视图漂移问题。
  • 在WorldArena和AgiBot-Challenge2026上分别排名第一、第二。
  • 适合研究机器人视觉、多视角生成与模型规划的学者使用。

世界基础模型(WFMs)虽具备强大模拟能力,但主要局限于单视角设置,缺乏机器人操作所需的多视角3D一致性。现有系统依赖多个摄像头(如第一人称、眼到手、腕装相机)进行策略学习,但多视图模型仅简单拼接视图标记,缺乏显式几何推理,导致跨视角物体漂移、深度不一致与纹理错位。我们发现根本原因在于缺少跨视角通信机制与3D几何先验。为此,提出PAIWorld框架,通过三个核心组件增强扩散-变压器世界模型:(1) 几何感知跨视角注意力块,建立视图间显式路径;(2) 几何旋转位置编码,将相机射线方向与外参融入注意力机制;(3) 潜在3D-REPA,从冻结的3D基础模型中蒸馏3D感知特征以保障3D一致性。基于DiT架构的PAIWorld在机器人操作基准测试中达到最优表现,在WorldArena榜单排名第一,在AgiBot-Challenge2026榜单排名第二,并支持模型基规划、世界动作模型及多视角策略后训练等下游应用。

原文摘要 · Abstract (English)

World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipulation. While robotic systems rely on multiple cameras (egocentric, eye-to-hand, and wrist-mounted) for policy learning, current multi-view world models simply concatenate view tokens without explicit geometric reasoning. This causes cross-view object drift, depth inconsistency, and texture misalignment. We trace these failures to two deficiencies: the absence of an explicit inter-view communication mechanism and the lack of a 3D geometric prior. We argue that resolving both simultaneously is necessary and sufficient. To address this, we present PAIWorld, a framework that augments diffusion-transformer world models via three core components: (1) Geometry-Aware Cross-View Attention blocks that establish an explicit pathway across views, (2) Geometric Rotary Position Embedding that encodes camera ray directions and extrinsic poses into the attention mechanism, and (3) Latent 3D-REPA, which distills 3D-aware features from frozen 3D foundation models to ensure 3D consistency. Built upon a DiT-based world foundation model, PAIWorld achieves state-of-the-art multi-view 3D consistency on robotic manipulation benchmarks, ranking 1st on the WorldArena leaderboard and 2nd on the AgiBot-Challenge2026 leaderboard, while enabling downstream applications such as model-based planning, world action models, and multi-view policy post-training.

3D一致性机器人操作扩散模型多视角建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。