用视觉问答生成城市风险地图,帮视障者安全导航
Urban Risk-Aware Navigation via VQA-Based Event Maps for People with Low Vision

- 通过多层提问框架让大模型理解复杂街景
- 在20城800图上实现四类风险分级,准确识别隐患
- 适合开发无障碍导航系统的研究者与开发者
全球数亿人存在视觉障碍,严重限制其在城市环境中独立安全出行。现有可穿戴辅助设备依赖专用视觉流水线,灵活性与泛化能力不足。本文提出基于视觉问答的事件地图框架,利用视觉语言模型(VLMs)在多样化真实环境中描述行人场景并识别危险,采用三层分层查询结构实现无需任务重训练的细粒度场景理解。模型输出经加权聚合形成风险评分系统,将街道段划分为四类安全等级,生成可用于路径规划的风险感知事件地图。为支持评估与后续研究,我们构建了一个覆盖六大洲20个城市的地理多样性数据集,包含超过800张标注图像和1.8万条问答。我们对四种VQA架构——ViLT、LLaVA、InstructBLIP和Qwen-VL进行基准测试,发现生成式多模态大模型显著优于分类方法,其中Qwen-VL在精度与召回率间表现最佳。结果表明,多模态大模型可作为视障人群辅助导航系统的灵活且通用基础。
原文摘要 · Abstract (English)
Visual impairment affects hundreds of millions of people worldwide, severely limiting their ability to navigate urban environments safely and independently. While wearable assistive devices offer a promising platform for real-time hazard detection, existing approaches rely on task-specific vision pipelines that lack flexibility and generalizability. In this work, we propose an event map framework based on visual question answering that leverages Vision-Language Models (VLMs) for pedestrian scene description and hazard identification across diverse real-world environments, using a three-level hierarchical query structure to enable fine-grained scene understanding without task-specific retraining. Model responses are aggregated into a weighted risk scoring system that maps street segments into four discrete safety categories, producing navigable risk-aware event maps for route planning. To support evaluation and future research, we introduce a geographically diverse dataset spanning 20 cities across six continents, comprising over 800 annotated images and 18,000 answered questions. We benchmark four VQA architectures -ViLT, LLaVA, InstructBLIP, and Qwen-VL- and find that generative Multimodal Large Language Models (MLLMs) substantially outperform classification-based approaches, with Qwen-VL achieving the best overall balance of precision and recall. These results demonstrate the viability of MLLMs as a flexible and generalizable foundation for assistive navigation systems for visually impaired people.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。