解耦语义与几何,提升机器人抓取的鲁棒性
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
- 分离语义与几何表示,融合噪声深度与语义先验
- 在RLBench上达87.4%成功率,噪声下提升13.9%-19.4%
- 适合需要低级空间线索和少样本泛化的场景
机器人操作需精准的空间理解以与真实世界物体交互。基于点的方法因采样稀疏而丢失细粒度语义;基于图像的方法通常将RGB与深度输入2D主干网络,该网络在3D辅助任务上预训练,但其语义与几何耦合,易受真实世界深度噪声干扰,破坏语义理解。此外,这些方法侧重高层几何,忽视对精确交互至关重要的低级空间线索。我们提出SpatialActor,一种解耦框架,显式分离语义与几何。语义引导的几何模块自适应融合来自噪声深度和语义引导专家先验的两种互补几何信息。同时,空间变换器利用低级空间线索实现准确的2D-3D映射,并促进空间特征间的交互。我们在50多个仿真与真实世界任务中评估了SpatialActor。其在RLBench上达到87.4%的性能,且在不同噪声条件下提升13.9%至19.4%,展现出强鲁棒性。此外,它显著提升了新任务的少样本泛化能力,并在多种空间扰动下保持稳健。
原文摘要 · Abstract (English)
Robotic manipulation requires precise spatial understanding to interact with objects in the real world. Point-based methods suffer from sparse sampling, leading to the loss of fine-grained semantics. Image-based methods typically feed RGB and depth into 2D backbones pre-trained on 3D auxiliary tasks, but their entangled semantics and geometry are sensitive to inherent depth noise in real-world that disrupts semantic understanding. Moreover, these methods focus on high-level geometry while overlooking low-level spatial cues essential for precise interaction. We propose SpatialActor, a disentangled framework for robust robotic manipulation that explicitly decouples semantics and geometry. The Semantic-guided Geometric Module adaptively fuses two complementary geometry from noisy depth and semantic-guided expert priors. Also, a Spatial Transformer leverages low-level spatial cues for accurate 2D-3D mapping and enables interaction among spatial features. We evaluate SpatialActor on multiple simulation and real-world scenarios across 50+ tasks. It achieves state-of-the-art performance with 87.4% on RLBench and improves by 13.9% to 19.4% under varying noisy conditions, showing strong robustness. Moreover, it significantly enhances few-shot generalization to new tasks and maintains robustness under various spatial perturbations. Project Page: https://shihao1895.github.io/SpatialActor
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。