arXiv:2508.12916cs.RO2025-08被引 1

仅用单个摄像头实现复杂环境下的智能物品抓取

RoboRetriever: Single-Camera Robot Object Retrieval via Active and Interactive Perception with Dynamic Scene Graph

  • 构建动态场景图,融合语义、几何与物体关系
  • 通过主动视角调整和交互感知提升识别准确率
  • 支持自然语言指令,适合真实家庭/办公场景

人类在杂乱、部分可见的环境中能凭借视觉推理、主动调整视角和物理交互轻松取物,仅依赖一对眼睛。而现有多数机器人系统需固定或多个摄像头配合完整视野,限制适应性且成本高。本文提出RoboRetriever,一种仅使用单个腕戴RGB-D相机和自然语言指令的现实世界物品检索框架。该系统将视觉观测与语义关联,构建并持续更新动态分层场景图,记录物体语义、几何信息及相互关系。监督模块基于此记忆与任务指令推断目标物体,并协调整合主动感知、交互感知与操作模块。为实现任务感知的场景对齐主动感知,引入新型视觉提示方法,利用大模型生成与语义目标和场景几何一致的6-DoF相机位姿。在包含人为干预的多样化真实场景中评估,RoboRetriever展现出强适应性和鲁棒性,仅用一个摄像头即可完成复杂物品检索任务。

原文摘要 · Abstract (English)

Humans effortlessly retrieve objects in cluttered, partially observable environments by combining visual reasoning, active viewpoint adjustment, and physical interaction-with only a single pair of eyes. In contrast, most existing robotic systems rely on carefully positioned fixed or multi-camera setups with complete scene visibility, which limits adaptability and incurs high hardware costs. We present \textbf{RoboRetriever}, a novel framework for real-world object retrieval that operates using only a \textbf{single} wrist-mounted RGB-D camera and free-form natural language instructions. RoboRetriever grounds visual observations to build and update a \textbf{dynamic hierarchical scene graph} that encodes object semantics, geometry, and inter-object relations over time. The supervisor module reasons over this memory and task instruction to infer the target object and coordinate an integrated action module combining \textbf{active perception}, \textbf{interactive perception}, and \textbf{manipulation}. To enable task-aware scene-grounded active perception, we introduce a novel visual prompting scheme that leverages large reasoning vision-language models to determine 6-DoF camera poses aligned with the semantic task goal and geometry scene context. We evaluate RoboRetriever on diverse real-world object retrieval tasks, including scenarios with human intervention, demonstrating strong adaptability and robustness in cluttered scenes with only one RGB-D camera.

机器人视觉感知自然语言动态场景图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。