arXiv:2503.08481cs.ROcs.CV2025-03CVPR被引 31

让视觉语言模型理解机器人能触达的范围,提升真实任务执行能力。

PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability

  • 用统一的空间可达性地图抽象不同机器人的物理能力
  • 在多机器人环境下比GPT-4o高14%准确率,超越多个先进模型
  • 适合机器人感知与具身智能研究者使用

理解环境和机器人物理可达性对任务执行至关重要。现有视觉语言模型(VLMs)在环境感知上表现优异,但在具身视觉推理任务中常因缺乏对机器人可达性的理解而生成不准确或不切实际的回答。为此,本文提出一种跨多种机器人的统一物理可达性表示——空间-物理可达性地图(S-P Map),并构建了集成该信息的视觉语言模型PhysVLM。S-P Map将机器人物理可达性抽象为与具体配置无关的通用空间表征,使模型聚焦于可达性特征而非机器人参数。PhysVLM通过新增特征编码器处理S-P Map,在不牺牲原有视觉语言能力的前提下实现物理可达性推理。为训练与评估,我们构建了大规模多机器人数据集Phys100K及挑战性基准EQA-phys,涵盖六种机器人在仿真与真实环境中的任务。实验表明,PhysVLM在EQA-phys上较GPT-4o提升14%,优于RoboMamba、SpatialVLM等先进模型,在RoboVQA-val和OpenEQA上也表现更优。S-P Map具有强兼容性,集成至GPT-4o-mini可带来7.1%性能提升。

原文摘要 · Abstract (English)

Understanding the environment and a robot's physical reachability is crucial for task execution. While state-of-the-art vision-language models (VLMs) excel in environmental perception, they often generate inaccurate or impractical responses in embodied visual reasoning tasks due to a lack of understanding of robotic physical reachability. To address this issue, we propose a unified representation of physical reachability across diverse robots, i.e., Space-Physical Reachability Map (S-P Map), and PhysVLM, a vision-language model that integrates this reachability information into visual reasoning. Specifically, the S-P Map abstracts a robot's physical reachability into a generalized spatial representation, independent of specific robot configurations, allowing the model to focus on reachability features rather than robot-specific parameters. Subsequently, PhysVLM extends traditional VLM architectures by incorporating an additional feature encoder to process the S-P Map, enabling the model to reason about physical reachability without compromising its general vision-language capabilities. To train and evaluate PhysVLM, we constructed a large-scale multi-robot dataset, Phys100K, and a challenging benchmark, EQA-phys, which includes tasks for six different robots in both simulated and real-world environments. Experimental results demonstrate that PhysVLM outperforms existing models, achieving a 14\% improvement over GPT-4o on EQA-phys and surpassing advanced embodied VLMs such as RoboMamba and SpatialVLM on the RoboVQA-val and OpenEQA benchmarks. Additionally, the S-P Map shows strong compatibility with various VLMs, and its integration into GPT-4o-mini yields a 7.1\% performance improvement.

机器人视觉语言模型可达性具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。