用预测能力减少误报,让视障者更安全地行走
ForeSightGuide: An Anticipatory Framework toward Accurate and Low-Redundancy Guidance for the Visually Impaired

- 结合视觉语言模型预测障碍物动态,提前识别危险
- 每条引导仅0.299条冗余警报,漏报率低至0.112
- 适合开发智能导盲系统或研究无障碍导航的团队
电子导行设备对视障人士独立出行至关重要。尽管视觉语言模型(VLMs)能提供丰富的环境理解,但在动态场景中常产生过多误报,导致认知过载。为此,我们提出ForeSightGuide——一种前瞻式辅助引导框架,将语义场景理解与预测性风险评估相结合。不同于传统反应式系统,ForeSightGuide利用VLM的推理能力预判障碍物运动,有效过滤非威胁对象,提供简洁、可操作的引导信息。为验证方法,我们构建了一个复杂动态真实交通场景的新数据集,用于评测预测性能。在公开基准和自建数据集上的大量实验表明,ForeSightGuide达到当前最优效果:每条引导输出仅产生0.299条冗余警报,同时保持0.112的低漏报率,证明其在安全导行中的有效性。
原文摘要 · Abstract (English)
Electronic travel aids are pivotal for the independent mobility of the visually impaired. While Vision-Language Models (VLMs) offer rich environmental understanding, they often suffer from excessive false positives in dynamic scenarios, leading to cognitive overload. To address this, we present ForeSightGuide, an anticipatory assistive guidance framework that couples semantic scene understanding with predictive hazard assessment. Unlike reactive systems, ForeSightGuide leverages the reasoning capabilities of VLMs to anticipate obstacle motion, effectively filtering out non-threatening objects to provide concise, actionable guidance. To validate our approach, we introduce a novel dataset captured in complex, dynamic real-world traffic scenes, designed to benchmark predictive capabilities. Extensive experiments on both public benchmarks and our proposed dataset demonstrate that ForeSightGuide achieves state-of-the-art performance. Notably, it significantly mitigates information overload by reducing redundant alerts to 0.299 per guidance output while maintaining a low missed-hazard rate of 0.112, proving its efficacy for safe walking assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。