用结构化物体槽表示提升机器人抓取成功率,不靠增加模型容量。
More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning

- 用物体中心槽注意力提取结构化特征,替代传统全局或密集局部特征。
- 在PickCube任务上达55.0%成功率,比基线高22.4%,无需微调编码器。
- 适合关注视觉-动作对齐、物体感知的机器人学习研究者。
机器人操作策略依赖预训练视觉模型,提供全局场景嵌入或密集补丁网格,二者混合了任务相关与无关特征。物体中心槽表示是一种结构化替代方案:将特征分组为少数每物体槽。我们在ManiSkill3 PickCube-v1上测试这种结构带来的收益,使用冻结编码器和保留种子评估。保持策略、目标令牌、渲染和校准不变,仅改变编码器,一个冻结的物体中心SPOT表示(DINO ViT-B/16 + 槽注意力)达到55.0±2.9%成功,比密集DINO全局特征基线(32.6±1.5%)高出22.4%,使用相同可训练策略且无编码器微调。单纯增加令牌数量无效:16倍密集补丁网格性能未提升。加入显式2D空间目标和原生分辨率渲染后,系统整体提升至68.7±4.2%,接近特权3D-Oracle上限(71.7±4.1%)。自动化运动学失败分类法区分空间精度(近失)与物体追踪(未抓取)失败:空间定位降低近失,但不影响未抓取;该分类法可迁移至更难的StackCube-v1,指向遮挡为主要瓶颈。
原文摘要 · Abstract (English)
Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid. Both mix task-relevant and task-irrelevant features. Object-centric slot representations are a structured alternative: they group features into a few per-object slots. We test what this structure buys on ManiSkill3 PickCube-v1, with a frozen encoder and a held-out-seed evaluation. Holding the policy, goal token, rendering, and calibration fixed and changing only the encoder, a frozen object-centric SPOT representation (DINO ViT-B/16 + Slot Attention) reaches 55.0$\pm$2.9% success, 22.4% above a dense DINO global-feature baseline (32.6 $\pm$ 1.5%), with the same trainable policy and no encoder fine-tuning. More tokens alone do not help: a dense patch grid with 16x the tokens performs no better than the global feature. Adding an explicit 2D spatial goal and native-resolution rendering raises the full system to 68.7$\pm$4.2%, just below a privileged 3D-oracle upper bound (71.7$\pm$4.1%). An automated kinematic failure taxonomy then separates spatial-precision (Near-Miss) failures from object-tracking (No-Grasp) failures: spatial grounding reduces Near-Miss while leaving No- Grasp unchanged. The same taxonomy transfers to the harder StackCube-v1 and points to occlusion as the main bottleneck.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。