构建2.2亿视障者可用的导航数据集,提升AI助盲能力
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
- 用人类+AI协作验证方式,高效生成真实场景图像描述
- 包含22000组图文对,818个问答样本测试深度感知能力
- 针对视障导航设计标准,适合做无障碍AI系统研究
全球超过22亿人受视力障碍影响,独立安全出行仍是重大挑战。尽管多模态大语言模型(MLLMs)为辅助导航带来新可能,但受限于缺乏可访问性意识的数据集,其发展受阻——因标注需大量专家投入。为此,我们推出GuideDog,一个包含22,000组图像-描述对(其中2,000组经人工验证)的真实世界行人场景数据集,覆盖46个国家。通过人类与AI协同的验证式标注流程,基于权威视障导航标准,显著提升可扩展性与质量。同时发布GuideDogQA,含818个样本的基准测试,评估物体识别与深度感知能力。实验表明,当前MLLM在深度感知及符合导航规范方面仍存在显著困难。
原文摘要 · Abstract (English)
For people affected by blindness and low vision (BLV), safe and independent navigation remains a major challenge, impacting over 2.2 billion individuals worldwide. Although multimodal large language models (MLLMs) offer new opportunities for assistive navigation, progress has been limited by the scarcity of accessibility-aware datasets, because creating them requires labor-intensive expert annotation. To this end, we introduce GuideDog, a novel dataset containing 22K image-description pairs (2K human-verified) capturing real-world pedestrian scenes across 46 countries. Our human-AI pipeline shifts annotation from generation to verification, grounded in established BLV guidance standards from experts and research, improving scalability while maintaining quality. We also present GuideDogQA, an 818-sample benchmark evaluating object recognition and depth perception. Experiments reveal that depth perception and adherence to these standards remain challenging for current MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。