为视障者开发轻量智能可穿戴系统,提升空间推理能力。
A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision
- 用多模态大模型微调增强空间理解能力
- 在VizWiz数据集上显著提升导航与识别准确率
- 轻量眼镜附件设计,适合日常使用
视障人士因缺乏视觉线索,在环境导航和物体定位方面面临重大挑战。空间推理对他们的独立行动至关重要。现有用于低视力人群的多模态大语言模型(MLLM)缺乏有效空间推理能力,且缺少轻量化、易用的系统支持。本文提出一种基于空间增强多模态大语言模型的新方法,通过微调使模型具备空间理解能力,显著提升对环境上下文的认知,助力导航与物体识别。硬件部分设计为眼镜附件,提高可及性与易用性。系统融合先进视觉语言模型,实时解析视觉信息并提供空间感知反馈。在VizWiz数据集上的评估显示准确率和用户体验均有显著提升,并构建了综合性真实场景数据集验证效果。
原文摘要 · Abstract (English)
People with blindness and low vision (pBLV) face significant challenges, struggling to navigate environments and locate objects due to limited visual cues. Spatial reasoning is crucial for these individuals, as it enables them to understand and interpret the spatial relationships in their surroundings, enhancing their ability to navigate and interact more safely and independently. Current multi-modal large language (MLLM) models for low vision people lack the spatial reasoning capabilities needed to effectively assist in these tasks. Moreover, there is a notable absence of lightweight, easy-to-use systems that allow pBLV to effectively perceive and interact with their surrounding environment. In this paper, we propose a novel spatial enhanced multi-modal large language model based approach for visually impaired individuals. By fine-tuning the MLLM to incorporate spatial reasoning capabilities, our method significantly improves the understanding of environmental context, which is critical for navigation and object recognition. The innovation extends to a hardware component, designed as an attachment for glasses, ensuring increased accessibility and ease of use. This integration leverages advanced VLMs to interpret visual data and provide real-time, spatially aware feedback to the user. Our approach aims to bridge the gap between advanced machine learning models and practical, user-friendly assistive devices, offering a robust solution for visually impaired users to navigate their surroundings more effectively and independently. The paper includes an in-depth evaluation using the VizWiz dataset, demonstrating substantial improvements in accuracy and user experience. Additionally, we design a comprehensive dataset to evaluate our method's effectiveness in realworld situations, demonstrating substantial improvements in accuracy and user experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。