arXiv:2510.11520cs.CV2025-10NeurIPS被引 4

构建多模态步行辅助数据集,提升视障者复杂环境导航能力

mmWalk: Towards Multi-modal Multi-view Walking Assistance

  • 构建含62000帧的多视角多模态步行数据集
  • 包含超55.9万张全景图像,覆盖视觉/深度/语义模态
  • 专为视障用户设计,适合研究无障碍导航与视觉问答

针对视障或低视力人群在极端或复杂环境中行走辅助仍面临挑战的问题,我们构建了mmWalk——一个模拟的多模态数据集,融合多视角传感器信息与适配无障碍需求的特征,用于户外安全导航。数据集包含120条人工控制、场景分类的步行轨迹,共62,000帧同步图像,涵盖超过559,000张全景图像,覆盖RGB、深度和语义模态。每条轨迹均包含真实场景中的边缘案例与视障用户相关的无障碍地标。此外,我们还构建了mmWalkVQA,一个包含超过69,000个视觉问答三元组的基准测试,覆盖9个类别,专为安全行走辅助设计。我们在零样本与少样本设置下评估主流视觉语言模型,发现其在风险评估与导航任务中表现不佳。通过在真实数据集上验证微调后的mmWalk模型,证明该数据集对推进多模态步行辅助的有效性。

原文摘要 · Abstract (English)

Walking assistance in extreme or complex environments remains a significant challenge for people with blindness or low vision (BLV), largely due to the lack of a holistic scene understanding. Motivated by the real-world needs of the BLV community, we build mmWalk, a simulated multi-modal dataset that integrates multi-view sensor and accessibility-oriented features for outdoor safe navigation. Our dataset comprises 120 manually controlled, scenario-categorized walking trajectories with 62k synchronized frames. It contains over 559k panoramic images across RGB, depth, and semantic modalities. Furthermore, to emphasize real-world relevance, each trajectory involves outdoor corner cases and accessibility-specific landmarks for BLV users. Additionally, we generate mmWalkVQA, a VQA benchmark with over 69k visual question-answer triplets across 9 categories tailored for safe and informed walking assistance. We evaluate state-of-the-art Vision-Language Models (VLMs) using zero- and few-shot settings and found they struggle with our risk assessment and navigational tasks. We validate our mmWalk-finetuned model on real-world datasets and show the effectiveness of our dataset for advancing multi-modal walking assistance.

视障辅助多模态导航系统数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。