arXiv:2409.17641cs.RO2024-09被引 4

用视觉语言模型指导机器人动态选视角,更好识别遮挡物体。

Scene Exploration by Vision-Language Models

  • 结合3D网格与视角调整,让机器人智能选择最佳观察位置。
  • 在遮挡和倾斜场景下,识别准确率显著高于传统方法。
  • 适合需要理解复杂环境的机器人应用,如家务助手。

主动感知使机器人通过动态调整视角来获取信息,是应对复杂、部分可观测环境的关键能力。本文提出AP-VLM框架,将主动感知与视觉语言模型(VLM)结合,指导机器人探索并回答语义查询。通过在场景上叠加3D虚拟网格并调整朝向,AP-VLM使机械臂能智能选择最优视角与朝向,解决遮挡或倾斜位置物体识别等难题。我们在两个平台(7-DOF Franka Panda和6-DOF UR5)上评估了系统,在不同物体配置的多个场景中测试。结果表明,AP-VLM显著优于被动感知方法和基线模型(如TGCSR),尤其在固定摄像头视角不适用时表现突出。该系统在真实场景中的适应性展示了其提升机器人对复杂环境理解的潜力,弥合了高层语义推理与底层控制之间的差距。

原文摘要 · Abstract (English)

Active perception enables robots to dynamically gather information by adjusting their viewpoints, a crucial capability for interacting with complex, partially observable environments. In this paper, we present AP-VLM, a novel framework that combines active perception with a Vision-Language Model (VLM) to guide robotic exploration and answer semantic queries. Using a 3D virtual grid overlaid on the scene and orientation adjustments, AP-VLM allows a robotic manipulator to intelligently select optimal viewpoints and orientations to resolve challenging tasks, such as identifying objects in occluded or inclined positions. We evaluate our system on two robotic platforms: a 7-DOF Franka Panda and a 6-DOF UR5, across various scenes with differing object configurations. Our results demonstrate that AP-VLM significantly outperforms passive perception methods and baseline models, including Toward Grounded Common Sense Reasoning (TGCSR), particularly in scenarios where fixed camera views are inadequate. The adaptability of AP-VLM in real-world settings shows promise for enhancing robotic systems' understanding of complex environments, bridging the gap between high-level semantic reasoning and low-level control.

机器人视觉语言模型主动感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。