arXiv:2603.15624cs.CVcs.AI2026-03被引 2

用视觉语言模型帮视障者导航,发现主流模型表现差异大。

Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision

  • 对比多个闭源与开源模型在视觉推理任务中的表现
  • GPT-4o 在空间理解与场景分析上显著优于其他模型
  • 开放模型在复杂环境适应性与推理精度上仍不足

本文研究视觉语言模型(VLMs)在帮助盲人和低视力人群(pBLV)进行导航任务中的潜力。评估了包括GPT-4V、GPT-4o、Gemini-1.5-Pro和Claude-3.5-Sonnet在内的闭源模型,以及Llava-v1.6-mistral和Llava-onevision-qwen等开源模型,分析其在基础视觉能力上的表现:计数周围障碍物、相对空间推理及与路径规划相关的常识性场景理解。进一步通过针对pBLV设计的特定提示,在导航场景中测试模型表现。结果表明,各模型间性能差异明显:GPT-4o在所有任务中均表现最优,尤其在空间推理与场景理解方面。而开源模型在复杂环境中的细微推理与适应能力较弱。常见问题包括在杂乱环境中误计物体数量、空间推理存在偏差,以及过度关注物体细节而忽略空间反馈,限制了其在导航辅助中的实用性。尽管如此,若能更好对齐人类反馈并提升空间推理能力,VLMs在导引辅助中仍具前景。本研究为开发者提供了当前VLMs优劣的实用洞见,指导其在辅助技术中有效集成,并解决关键局限以提升可用性。

原文摘要 · Abstract (English)

This paper investigates the potential of vision-language models (VLMs) to assist people with blindness and low vision (pBLV) in navigation tasks. We evaluate state-of-the-art closed-source models, including GPT-4V, GPT-4o, Gemini-1.5-Pro, and Claude-3.5-Sonnet, alongside open-source models, such as Llava-v1.6-mistral and Llava-onevision-qwen, to analyze their capabilities in foundational visual skills: counting ambient obstacles, relative spatial reasoning, and common-sense wayfinding-pertinent scene understanding. We further assess their performance in navigation scenarios, using pBLV-specific prompts designed to simulate real-world assistance tasks. Our findings reveal notable performance disparities between these models: GPT-4o consistently outperforms others across all tasks, particularly in spatial reasoning and scene understanding. In contrast, open-source models struggle with nuanced reasoning and adaptability in complex environments. Common challenges include difficulties in accurately counting objects in cluttered settings, biases in spatial reasoning, and a tendency to prioritize object details over spatial feedback, limiting their usability for pBLV in navigation tasks. Despite these limitations, VLMs show promise for wayfinding assistance when better aligned with human feedback and equipped with improved spatial reasoning. This research provides actionable insights into the strengths and limitations of current VLMs, guiding developers on effectively integrating VLMs into assistive technologies while addressing key limitations for enhanced usability.

视觉语言模型导航辅助视障支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。