提升双臂机器人多视角感知,让视觉理解更准、空间判断更稳。
MV-Actor: Aligning Multi-View Semantics and Spatial Awareness for Bimanual Manipulation

- 通过多视角语义交互共享跨视图理解
- 在仿真中达87.8%成功率,超越现有方法
- 适合需要高精度空间感知的双臂操作场景
机器人操作已在工业场景广泛应用。相比单臂操作,双臂操作配备多摄像头以获取不同视角信息。然而,现有多视角策略通常独立编码各视角或浅层融合特征,导致语义感知共享不足且空间感知不可靠。本文提出MV-Actor框架,构建双臂操作的统一语义-空间表征:首先通过多视角语义交互实现跨视图语义共享;再利用语义-空间标记交互,将视觉语义与前馈重建模型特征结合,获得可靠的空间感知;最后引入引导式度量深度修复模块,优化消费级深度传感器噪声下的深度数据,提供更可靠的度量基准。在PerAct2双臂基准上的仿真实验表明,MV-Actor达到87.8%的平均成功率,处于领先水平。在真实世界中面对频繁视角变化和不稳定深度数据的情况下,其性能优于仅用RGB和RGB-D的基线方法,验证了共享语义感知与可靠空间意识对双臂操作的关键价值。
原文摘要 · Abstract (English)
Robotic manipulation has been widely applied in industrial scenarios. Compared with single-arm manipulation, bimanual manipulation is equipped with multiple cameras to capture information from different viewpoints. However, existing multi-view policies encode each view independently or fuse view features shallowly, resulting in limited sharing semantic perception and unreliable spatial awareness. In this paper, we propose \textbf{MV-Actor}, a multi-view perception framework that builds a unified semantic-spatial representation for bimanual manipulation. First, MV-Actor performs Multi-view Semantic Interaction to share semantic perception across views. Then it uses Semantic-Spatial Token Interaction to ground visual semantics with feed-forward reconstruction model features and acquire reliable spatial awareness. Finally, a Guided Metric Depth Repair module refines degraded sensor depth to provide more reliable metric anchors under consumer-grade depth noise. In simulation experiments conducted on the PerAct2 bimanual benchmark, MV-Actor achieves a state-of-the-art average success rate of 87.8\%. In real-world evaluations with more frequent viewpoint changes and unstable consumer-grade depth, MV-Actor outperforms both RGB and RGB-D baselines, further demonstrating the benefit of sharing semantic perception and reliable spatial awareness for bimanual manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。