用对象中心表示法提升机器人抓取的泛化能力
Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation
- 将密集特征分组为有限对象实体,减少无关信息干扰
- 在光照、纹理变化下仍保持高泛化性能,优于全局与密集特征
- 无需任务预训练,适合真实动态环境中的机器人视觉系统
机器人抓取策略的泛化能力高度依赖于视觉表征的选择。现有方法通常使用预训练编码器提取两类特征:全局特征(通过池化得到单一向量)和密集特征(保留最后一层的局部嵌入)。这两类特征均混杂任务相关与无关信息,在光照、纹理变化或存在干扰物等分布外场景下表现不佳。本文提出一种中间结构化的替代方案:基于槽位的对象中心表征(SBOCR),将密集特征聚类为有限数量的对象类实体。该表征能自然抑制噪声输入,同时保留完成任务所需的关键信息。我们在一系列模拟与真实世界的抓取任务中对比了多种全局、密集及槽位表征,评估其在光照、纹理变化及干扰物存在下的泛化能力。结果表明,SBOCR驱动的策略在无任务特定预训练条件下,仍显著优于基于全局与密集特征的策略,验证了其在动态真实环境中的有效性。
原文摘要 · Abstract (English)
The generalization capabilities of robotic manipulation policies are heavily influenced by the choice of visual representations. Existing approaches typically rely on representations extracted from pre-trained encoders, using two dominant types of features: global features, which summarize an entire image via a single pooled vector, and dense features, which preserve a patch-wise embedding from the final encoder layer. While widely used, both feature types mix task-relevant and irrelevant information, leading to poor generalization under distribution shifts, such as changes in lighting, textures, or the presence of distractors. In this work, we explore an intermediate structured alternative: Slot-Based Object-Centric Representations (SBOCR), which group dense features into a finite set of object-like entities. This representation permits to naturally reduce the noise provided to the robotic manipulation policy while keeping enough information to efficiently perform the task. We benchmark a range of global and dense representations against intermediate slot-based representations, across a suite of simulated and real-world manipulation tasks ranging from simple to complex. We evaluate their generalization under diverse visual conditions, including changes in lighting, texture, and the presence of distractors. Our findings reveal that SBOCR-based policies outperform dense and global representation-based policies in generalization settings, even without task-specific pretraining. These insights suggest that SBOCR is a promising direction for designing visual systems that generalize effectively in dynamic, real-world robotic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。