arXiv:2412.20059cs.CV2024-12被引 18

用AI眼镜实时识别物体并语音描述环境,帮视障者安全独立生活

AI-based Wearable Vision Assistance System for the Visually Impaired: Integrating Real-Time Object Recognition and Contextual Understanding Using Large Vision-Language Models

  • 基于大视觉语言模型,通过一键添加新对象提升识别准确率
  • 集成距离传感器,碰撞前用蜂鸣声提醒,保障行走安全
  • 硬件轻便(树莓派4+帽子摄像头),适合日常使用,适合视障人群

视觉障碍影响人们过上正常生活的能力,他们在阅读、书写、出行和社交活动中面临困难。传统辅助手段难以获取丰富上下文环境信息。本文提出一种新型可穿戴视觉辅助系统,采用帽子搭载摄像头与树莓派4模型B(8GB内存)结合,利用人工智能技术实时通过声音提示反馈用户。系统支持一键添加新人物或物体数据,持续提升识别精度;借助大视觉语言模型(LVLM)对环境物体提供详细语音描述;同时配备距离传感器,当用户即将碰撞障碍物时,蜂鸣器会立即发出警报,确保导航安全。全面评估表明,该系统融合硬件与AI(含LVLM与物联网)的创新设计,在辅助技术方面实现显著进步,有效解决视障群体的核心需求。

原文摘要 · Abstract (English)

Visual impairment affects the ability of people to live a life like normal people. Such people face challenges in performing activities of daily living, such as reading, writing, traveling and participating in social gatherings. Many traditional approaches are available to help visually impaired people; however, these are limited in obtaining contextually rich environmental information necessary for independent living. In order to overcome this limitation, this paper introduces a novel wearable vision assistance system that has a hat-mounted camera connected to a Raspberry Pi 4 Model B (8GB RAM) with artificial intelligence (AI) technology to deliver real-time feedback to a user through a sound beep mechanism. The key features of this system include a user-friendly procedure for the recognition of new people or objects through a one-click process that allows users to add data on new individuals and objects for later detection, enhancing the accuracy of the recognition over time. The system provides detailed descriptions of objects in the user's environment using a large vision language model (LVLM). In addition, it incorporates a distance sensor that activates a beeping sound using a buzzer as soon as the user is about to collide with an object, helping to ensure safety while navigating their environment. A comprehensive evaluation is carried out to evaluate the proposed AI-based solution against traditional support techniques. Comparative analysis shows that the proposed solution with its innovative combination of hardware and AI (including LVLMs with IoT), is a significant advancement in assistive technology that aims to solve the major issues faced by the community of visually impaired people

视觉辅助AI助残大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。