arXiv:2606.26800cs.RO2026-06中稿 · 2026 IEEE/RSJ Inte…被引 1

用纯视觉构建结构化场景接口,让机器人少样本学会复杂操作。

SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation

论文配图:SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation
图 1 · 摘自论文原文
  • 设计统一的RGB-only场景接口,融合深度、物体布局和运动轨迹
  • 10次演示下比最强基线提升15%,接近50次演示效果
  • 适合少样本、跨设备、需空间推理的现实机器人任务

真实世界机器人操作需要空间定位、任务感知和精确控制,但在数据稀缺情况下学习尤为困难。现有方法常在可扩展的任务推理与显式物理结构间权衡:视频方法长期会几何漂移,3D方法依赖深度传感,多数流/轨迹接口缺乏显式RGB几何表示。本文提出SSI-Policy,基于结构化场景接口(SSI)——一种统一的纯RGB中间表示,联合编码单目深度特征、语言引导的物体布局及指令条件的2D运动轨迹。关键在于,SSI与机器人无关,且可从无动作视频中训练,将感知与控制解耦,使下游策略能从少量示范学习。在只需每任务10次示范的LIBERO基准上,该方法较最强基线提升近15%,表现媲美需大规模外部预训练的50次示范方法。消融实验表明几何与运动线索在共享接口中提供互补优势。进一步在13个真实任务上验证,涵盖空间推理、跨设备迁移与接触密集操作。

原文摘要 · Abstract (English)

Real-world robotic manipulation demands spatial grounding, task-aware reasoning, and precise control. Learning such capabilities becomes particularly challenging in the low-data regime. Prior methods often trade off scalable task-level reasoning and explicit physical structure: video-based approaches can drift geometrically over long horizons, 3D approaches often require depth sensing, and many flow/trajectory interfaces emphasize motion without an explicit RGB-only geometric representation. We introduce SSI-Policy, a modular framework built around a Structured Scene Interface (SSI) -- a unified, RGB-only intermediate representation that jointly encodes monocular depth features, language-grounded object layouts, and instruction-conditioned 2D motion trajectories. Critically, SSI is robot-agnostic and trainable from action-free video, decoupling perception from control so that the downstream policy can learn from few demonstrations. On the LIBERO benchmark with only 10 demonstrations per task, SSI-Policy improves over the strongest prior method by nearly 15\% and remains competitive with 50-demo methods that leverage large-scale external pretraining. Ablations show that geometric and motion cues provide complementary benefits within the shared interface. We further validate on 13 real-world tasks spanning spatial reasoning, cross-embodiment transfer, and contact-rich manipulation.

机器人操作视觉语言少样本学习结构化表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。