arXiv:2511.10615cs.CVcs.CL2025-11被引 3

轻量级视觉语言模型助力视障用户获取更精准的环境描述。

Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals

  • 设计双评估框架,分别聚焦空间关系与导航辅助信息。
  • 500M和2.2B参数模型在户外/室内数据集上均表现良好。
  • 适配手机端运行,支持低精度推理,适合资源受限设备。

大型视觉语言模型(VLMs)虽能出色理解与生成视频描述,但其高内存、计算与部署需求限制了实际应用,尤其对依赖详细、上下文感知描述的盲人及低视力(BLV)用户而言。为研究模型规模对可访问性描述质量的影响,我们在两个不同数据集(AVCaps:户外;Charades:室内)上评估了具有500M和2.2B参数的SmolVLM2变体。本文提出两种专为BLV可访问性设计的新评估框架:多情境BLV框架(评估空间方位、社交互动、动作事件与氛围情境)和导航辅助框架(关注移动关键信息)。此外,系统评估了四种提示设计策略,并将两个模型部署于智能手机上,对比了FP32与INT8精度版本,以分析资源受限移动设备上的实际性能表现。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) excel at understanding and generating video descriptions but their high memory, computation, and deployment demands hinder practical use particularly for blind and low-vision (BLV) users who depend on detailed, context-aware descriptions. To study the effect of model size on accessibility-focused description quality, we evaluate SmolVLM2 variants with 500M and 2.2B parameters across two diverse datasets: AVCaps (outdoor), and Charades (indoor). In this work, we introduce two novel evaluation frameworks specifically designed for BLV accessibility assessment: the Multi-Context BLV Framework evaluating spatial orientation, social interaction, action events, and ambience contexts; and the Navigational Assistance Framework focusing on mobility-critical information. Additionally, we conduct a systematic evaluation of four different prompt design strategies and deploy both models on a smartphone, evaluating FP32 and INT8 precision variants to assess real-world performance constraints on resource-limited mobile devices.

视觉语言模型视障辅助轻量化部署移动端推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。