提出新模型SEM,让机器人更懂空间,操作更稳健。
SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
- 用3D几何上下文增强视觉表征,提升空间感知
- 通过图模型捕捉关节依赖关系,建模机器人本体特征
- 在多样任务中表现超越现有方法,适合复杂操作场景
机器人操作的核心挑战在于构建具备强空间理解能力的策略模型,即能够推理三维几何、物体关系与机器人本体特性。现有方法存在局限:3D点云模型缺乏语义抽象,2D图像编码器难以进行空间推理。为此,我们提出SEM(Spatial Enhanced Manipulation)模型,一种基于扩散的新型策略框架,从两个互补角度显式增强空间理解。空间增强模块为视觉表征注入3D几何上下文,机器人状态编码器则通过图结构建模关节依赖关系,捕捉本体感知特征。融合两者后,SEM显著提升了空间理解能力,在多种任务中实现更强鲁棒性与泛化性,性能优于现有基线方法。
原文摘要 · Abstract (English)
A key challenge in robot manipulation lies in developing policy models with strong spatial understanding, the ability to reason about 3D geometry, object relations, and robot embodiment. Existing methods often fall short: 3D point cloud models lack semantic abstraction, while 2D image encoders struggle with spatial reasoning. To address this, we propose SEM (Spatial Enhanced Manipulation model), a novel diffusion-based policy framework that explicitly enhances spatial understanding from two complementary perspectives. A spatial enhancer augments visual representations with 3D geometric context, while a robot state encoder captures embodiment-aware structure through graphbased modeling of joint dependencies. By integrating these modules, SEM significantly improves spatial understanding, leading to robust and generalizable manipulation across diverse tasks that outperform existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。