arXiv:2412.20903cs.CVcs.AI2024-12ICCV被引 18

用视觉语言模型帮视障者实时导航,首建大规模数据集与评测基准。

WalkVLM:Aid Visually Impaired People Walking by Vision Language Model

  • 构建12,000对视频-标注数据,首次提供步行辅助统一基准。
  • 提出分层思维链规划,生成简洁高效提醒,减少冗余信息。
  • 适合研究无障碍导航、多模态推理与实时系统优化的开发者。

全球约2亿人患有不同程度的视觉障碍,利用AI技术为他们提供行走辅助至关重要。近年来,视觉语言模型(VLMs)在该领域受到关注,但现有方法多依赖自建问答数据集,且未公开,缺乏标准化训练与评估基准。此外,行走辅助需实时分析视频流并生成简明准确的提示,而传统VLM存在响应过长、推理效率低的问题。本文提出首个专注于步行辅助的大规模数据集,包含12,000个视频-标注对,为训练与评估提供统一基准。同时设计了WalkVLM模型,采用思维链进行分层规划以生成简洁信息,结合时间感知自适应预测机制降低提醒冗余。实验表明,该模型在流式视频处理中显著优于其他VLM。相关数据集与代码已开源:https://walkvlm2024.github.io。

原文摘要 · Abstract (English)

Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people. With the recent progress of vision-language models (VLMs), applying VLMs to offer walking guidance has become popular. However, the existing methods of walking guidance are mainly based on self-curated question-answering datasets that are not publicly accessible, without a standardized benchmark for training or evaluation. Moreover, walking assistance often requires real-time streaming video analysis and the generation of concise yet informative reminders, making VLMs struggle due to excessive responses and low efficiency in inferences. In this paper, we introduce the first large-scale dataset dedicated to walking assistance, comprising 12,000 video-annotation pairs, to provide a unified benchmark for training and evaluating systems to help visually-impaired individuals walk. Furthermore, a WalkVLM model is proposed, which employs chain of thought for hierarchical planning to generate concise but informative reminders and utilizes temporal-aware adaptive prediction to reduce the temporal redundancy of reminders. Finally, we have established a solid benchmark for blind walking task and verified the advantages of WalkVLM in stream video processing for this task compared to other VLMs. Our dataset and code are available at https://walkvlm2024.github.io.

视障辅助视觉语言模型实时推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。