评估视觉语言模型在机器人运动空间推理中的表现,探索其用于规划具偏好动作的可行性。
Evaluating VLMs' Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences
- 用四种查询方法测试主流VLM对机器人运动空间关系的理解能力。
- 最优方法下Qwen2.5-VL零样本准确率达71.4%,微调后达75%。
- 揭示了准确率与计算开销(令牌数)之间的权衡,适合机器人规划研究者参考。
理解用户指令及周围环境中的物体空间关系,对智能机器人系统完成多样化任务至关重要。视觉语言模型(VLMs)的自然语言与空间推理能力有望提升机器人规划器在新任务、新物体和运动规范下的泛化能力。尽管基础模型已应用于任务规划,但其在执行用户偏好或约束(如物体距离、拓扑属性、运动风格)所需的时空推理能力仍不明确。本文评估了四种先进VLM在机器人运动空间推理上的表现,采用四种不同查询方法。结果表明,在最佳查询方式下,Qwen2.5-VL实现71.4%的零样本准确率,微调后达到75%;GPT-4o表现较低。我们评估了两类运动偏好(物体邻近性与路径风格),并分析了准确率与计算成本(令牌数)间的权衡。该工作展示了将VLM集成至机器人运动规划流程的潜力。
原文摘要 · Abstract (English)
Understanding user instructions and object spatial relations in surrounding environments is crucial for intelligent robot systems to assist humans in various tasks. The natural language and spatial reasoning capabilities of Vision-Language Models (VLMs) have the potential to enhance the generalization of robot planners on new tasks, objects, and motion specifications. While foundation models have been applied to task planning, it is still unclear the degree to which they have the capability of spatial reasoning required to enforce user preferences or constraints on motion, such as desired distances from objects, topological properties, or motion style preferences. In this paper, we evaluate the capability of four state-of-the-art VLMs at spatial reasoning over robot motion, using four different querying methods. Our results show that, with the highest-performing querying method, Qwen2.5-VL achieves 71.4% accuracy zero-shot and 75% on a smaller model after fine-tuning, and GPT-4o leads to lower performance. We evaluate two types of motion preferences (object-proximity and path-style), and we also analyze the trade-off between accuracy and computation cost in number of tokens. This work shows some promise in the potential of VLM integration with robot motion planning pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。