arXiv:2601.20381cs.RO2026-01被引 1

让机器人视觉模型学会聚焦关键物体,提升操作鲁棒性。

STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation

  • 用任务感知的槽结构增强冻结视觉模型,实现轻量级对象中心表征。
  • 在多种干扰条件下,操控性能和泛化能力显著优于传统方法。
  • 适合需要精准物体理解的机器人抓取与复杂场景操作任务。

视觉基础模型虽具备强大感知能力,但其密集表征缺乏显式的对象级结构,限制了机器人操作中的鲁棒性和可控性。本文提出 STORM(Slot-based Task-aware Object-centric Representation for robotic Manipulation),一种轻量级的对象中心适配模块,通过少量任务感知槽,为冻结的视觉基础模型注入对象中心表征。STORM 采用两阶段训练策略:首先在冻结主干网络上,利用语言嵌入进行视觉-语义预训练,训练少数几层对象中心表示;随后与下游操控策略联合优化,实现任务对齐。该策略避免了槽结构退化,保持语义一致性,同时使感知与任务目标对齐。在对象发现基准和机器人操作任务上的实验表明,相比直接使用冻结或微调后的基础模型特征,或现有对象中心表示,STORM 在视觉扰动(如干扰物、纹理、光照变化)下均展现出更优的控制性能与泛化能力。STORM 不仅是优化通用基础模型特征的高效机制,也为策略学习注入有益的结构与语义先验。

原文摘要 · Abstract (English)

Visual foundation models provide strong perceptual features for robotics, but their dense representations lack explicit object-level structure, limiting robustness and controllability in manipulation tasks. We propose STORM (Slot-based Task-aware Object-centric Representation for robotic Manipulation), a lightweight object-centric adaptation module that augments frozen visual foundation models with a small set of task-aware slots for robotic manipulation. Rather than fully tuning large backbones on the task, STORM employs an efficient two-stage training strategy: few layers of object-centric representation are first trained on top of the frozen backbone through visual--semantic pretraining using language embeddings, then jointly adapted with a downstream manipulation policy for task alignement. This staged learning prevents degenerate slot formation and preserves semantic consistency while aligning perception with task objectives. Experiments on object discovery benchmarks and robotic manipulation tasks show that STORM improves control performance and generalization to visual shifts (distractors, textures, lighting) compared to directly using frozen or fine-tuned foundation model features, or existing object-centric representations. STORM serves not only as an efficient mechanism for refining generic foundation model features, but also as a novel way of injecting beneficial structural and semantic bias into policy learning.

机器人操作对象中心视觉表征任务感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。