arXiv:2502.07183cs.ROcs.CV2025-02ICRA被引 6

提升视障导引机器人对空间关系的理解能力,让导航提示更精准。

Space-Aware Instruction Tuning: Dataset and Benchmark for Guide Dog Robots Assisting the Visually Impaired

  • 构建基于3D空间路径的指令微调数据集,强化模型对环境的空间感知。
  • 在复杂场景下,新模型的导航指导准确率显著优于现有算法。
  • 适合研究无障碍AI、具身智能与视觉语言模型的应用者参考。

导盲机器人有望提升视障人士的出行安全与独立性,克服传统导盲犬在感知智能与沟通上的局限。随着视觉语言模型(VLMs)的发展,机器人可生成周围环境的自然语言描述,辅助安全决策。然而,现有VLMs在理解与传达空间关系方面表现不佳,这在街口等复杂环境中尤为关键。为此,我们提出了空间感知指令微调数据集(SAIT)和空间感知基准(SA-Bench),通过自动化数据生成流程聚焦于3D空间中的虚拟路径及周边环境,增强模型对环境的整体理解,使VLM能提供更精确的行走指引。同时提出评估协议以衡量VLM在提供步行引导方面的效能。对比实验表明,我们的空间感知微调模型优于当前最优算法。SAIT数据集与SA-Bench已完全开源,代码可在 https://github.com/byungokhan/Space-awareVLM 获取。

原文摘要 · Abstract (English)

Guide dog robots offer promising solutions to enhance mobility and safety for visually impaired individuals, addressing the limitations of traditional guide dogs, particularly in perceptual intelligence and communication. With the emergence of Vision-Language Models (VLMs), robots are now capable of generating natural language descriptions of their surroundings, aiding in safer decision-making. However, existing VLMs often struggle to accurately interpret and convey spatial relationships, which is crucial for navigation in complex environments such as street crossings. We introduce the Space-Aware Instruction Tuning (SAIT) dataset and the Space-Aware Benchmark (SA-Bench) to address the limitations of current VLMs in understanding physical environments. Our automated data generation pipeline focuses on the virtual path to the destination in 3D space and the surroundings, enhancing environmental comprehension and enabling VLMs to provide more accurate guidance to visually impaired individuals. We also propose an evaluation protocol to assess VLM effectiveness in delivering walking guidance. Comparative experiments demonstrate that our space-aware instruction-tuned model outperforms state-of-the-art algorithms. We have fully open-sourced the SAIT dataset and SA-Bench, along with the related code, at https://github.com/byungokhan/Space-awareVLM

导盲机器人空间感知视觉语言模型无障碍AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。