arXiv:2507.12391cs.RO2025-07中稿 · the 2025 SICE Fest…被引 4

测试15个多模态大模型在机器人路径规划中视觉输入的价值

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning

  • 对比文本与图文输入在2D网格中的路径生成效果
  • 小网格下视觉或提示词提升成功率,大网格性能显著下降
  • 当前模型在空间推理和多模态融合上仍存明显短板

大型语言模型(LLMs)在提升机器人路径规划方面展现出潜力。本文通过全面基准测试,评估了多模态LLMs在该任务中视觉输入的效用。我们在2D网格环境(模拟简化机器人规划)中评估了15个多模态LLMs,比较了文本仅输入与文本+视觉输入在不同模型规模和网格复杂度下的表现。结果显示,在较简单的小型网格中,视觉输入或少量示例提示词带来一定优势,成功率为中等;但在更大网格中,性能显著下降,凸显可扩展性挑战。虽然更大模型平均成功率更高,但视觉模态并未在所有情况下超越结构良好文本。简单网格上的成功路径质量普遍较高。这些结果表明,当前模型在鲁棒空间推理、约束遵守和可扩展多模态整合方面存在局限,指明了未来在机器人路径规划中发展多模态大模型的关键方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show potential for enhancing robotic path planning. This paper assesses visual input's utility for multimodal LLMs in such tasks via a comprehensive benchmark. We evaluated 15 multimodal LLMs on generating valid and optimal paths in 2D grid environments, simulating simplified robotic planning, comparing text-only versus text-plus-visual inputs across varying model sizes and grid complexities. Our results indicate moderate success rates on simpler small grids, where visual input or few-shot text prompting offered some benefits. However, performance significantly degraded on larger grids, highlighting a scalability challenge. While larger models generally achieved higher average success, the visual modality was not universally dominant over well-structured text for these multimodal systems, and successful paths on simpler grids were generally of high quality. These results indicate current limitations in robust spatial reasoning, constraint adherence, and scalable multimodal integration, identifying areas for future LLM development in robotic path planning.

多模态路径规划大模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。