arXiv:2602.18374cs.ROcs.AI2026-02

让机器人零样本交互感知,靠视觉大模型+推拉动作解构复杂场景

Zero-shot Interactive Perception

  • 用推线增强视觉输入,结合大模型决策推、抓等动作
  • 在7自由度机械臂上测试,推动作成功率显著高于传统方法
  • 适合解决遮挡多、信息不全的现实操作任务

交互感知(IP)使机器人通过物理交互获取环境隐藏信息并执行操作计划,对处理复杂、部分可观测场景中的遮挡和模糊至关重要。本文提出零样本交互感知(ZS-IP)框架,将多策略操作(推与抓)与记忆驱动的视觉语言模型(VLM)结合,引导机器人交互并回答语义查询。ZS-IP包含三个核心组件:(1) 增强观测(EO)模块,通过常规关键点和新提出的推线(pushlines)——一种专为推动作设计的2D视觉增强,提升视觉感知;(2) 基于记忆的动作模块,通过上下文检索强化语义推理;(3) 机器人控制器,根据VLM输出执行推、拉或抓。不同于为拾取放置优化的网格增强,推线捕捉接触丰富动作的可操作性,显著提升推动作性能。我们在7自由度Franka Panda机械臂上,在多种含遮挡和任务复杂度的场景中评估了ZS-IP,结果表明其优于被动感知和视角感知方法(如基于标记的视觉提示MOKA),尤其在推动作上表现更优,同时保持非目标物体完整性。

原文摘要 · Abstract (English)

Interactive perception (IP) enables robots to extract hidden information in their workspace and execute manipulation plans by physically interacting with objects and altering the state of the environment -- crucial for resolving occlusions and ambiguity in complex, partially observable scenarios. We present Zero-Shot IP (ZS-IP), a novel framework that couples multi-strategy manipulation (pushing and grasping) with a memory-driven Vision Language Model (VLM) to guide robotic interactions and resolve semantic queries. ZS-IP integrates three key components: (1) an Enhanced Observation (EO) module that augments the VLM's visual perception with both conventional keypoints and our proposed pushlines -- a novel 2D visual augmentation tailored to pushing actions, (2) a memory-guided action module that reinforces semantic reasoning through context lookup, and (3) a robotic controller that executes pushing, pulling, or grasping based on VLM output. Unlike grid-based augmentations optimized for pick-and-place, pushlines capture affordances for contact-rich actions, substantially improving pushing performance. We evaluate ZS-IP on a 7-DOF Franka Panda arm across diverse scenes with varying occlusions and task complexities. Our experiments demonstrate that ZS-IP outperforms passive and viewpoint-based perception techniques such as Mark-Based Visual Prompting (MOKA), particularly in pushing tasks, while preserving the integrity of non-target elements.

交互感知视觉语言模型机器人操作零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。