用大模型+图像分割辅助视障者实时理解周围环境。
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
- 将图像分割结果作为外部知识输入大视觉语言模型,减少错误描述。
- 在POPE、MME等数据集上优于Qwen-VL-Chat,准确率提升显著。
- 支持长按、点击、双击操作,适配视障用户交互习惯。
由于自然场景复杂,视障人士感知周围环境面临巨大挑战,个人与社会活动受限严重。本文提出一种基于大视觉语言模型(LVLM)的环境感知系统,通过可穿戴设备捕捉当前场景,并让用户通过屏幕操作获取分析结果。用户长按屏幕可获全局场景描述,轻点或滑动可查询场景中物体类别(由分割模型生成),双击则可获取感兴趣物体的详细描述。为提升感知准确性,本文创新性地将RGB图像的分割结果作为外部知识输入LVLM,有效降低模型幻觉。技术实验表明,该系统在POPE、MME和LLaVA-QA90数据集上优于Qwen-VL-Chat;探索性实验显示,系统能有效帮助视障人士理解周围环境。
原文摘要 · Abstract (English)
It is a challenging task for visually impaired people to perceive their surrounding environment due to the complexity of the natural scenes. Their personal and social activities are thus highly limited. This paper introduces a Large Vision-Language Model(LVLM) based environment perception system which helps them to better understand the surrounding environment, by capturing the current scene they face with a wearable device, and then letting them retrieve the analysis results through the device. The visually impaired people could acquire a global description of the scene by long pressing the screen to activate the LVLM output, retrieve the categories of the objects in the scene resulting from a segmentation model by tapping or swiping the screen, and get a detailed description of the objects they are interested in by double-tapping the screen. To help visually impaired people more accurately perceive the world, this paper proposes incorporating the segmentation result of the RGB image as external knowledge into the input of LVLM to reduce the LVLM's hallucination. Technical experiments on POPE, MME and LLaVA-QA90 show that the system could provide a more accurate description of the scene compared to Qwen-VL-Chat, exploratory experiments show that the system helps visually impaired people to perceive the surrounding environment effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。