为视障人士设计更实用的AI视觉描述评估框架
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
- 构建视障用户偏好数据集,模拟真实导航需求
- 新评估器在人类一致性与效率上均优于现有方法
- 适合研究无障碍AI辅助系统或人机交互的开发者
大型视觉语言模型(LVLMs)在辅助视障或低视力(BLV)人群方面展现出巨大潜力。然而,衡量其在真实场景中的实际效用极具挑战性,因为评估描述是否对BLV用户有用,需区别于常规场景描述的评价方式。尽管“视觉语言模型作为评估指标”范式已出现,但现有评估器仍无法满足BLV导向评估的关键需求:(1)与人类判断高度相关;(2)理解长指令;(3)评分生成高效;(4)多维度评估。为此,我们提出统一框架,弥合自动化评估与真实BLV需求之间的差距。首先,通过面向BLV用户的深入研究,量化其导航偏好,构建大规模用户模拟偏好数据集VL-GUIDEDATA,包含图像-请求-响应-评分对。随后,基于该数据集开发出具备可访问性感知能力的评估器VL-GUIDE-S,其在人类对齐度和推理效率上均超越现有(大)视觉语言模型裁判。尤为关键的是,其有效性不仅限于单一领域,在多个细粒度、对BLV至关重要的维度上表现优异。我们希望本工作能为安全、无障碍导航的自动AI裁判奠定基础。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) demonstrate a promising direction for assisting individuals with blindness or low-vision (BLV). Yet, measuring their true utility in real-world scenarios is challenging because evaluating whether their descriptions are BLV-informative requires a fundamentally different approach from assessing standard scene descriptions. While the "VLM-as-a-metric" or "LVLM-as-a-judge" paradigm has emerged, existing evaluators still fall short of capturing the unique requirements of BLV-centric evaluation, lacking at least one of the following key properties: (1) High correlation with human judgments, (2) Long instruction understanding, (3) Score generation efficiency, and (4) Multi-dimensional assessment. To this end, we propose a unified framework to bridge the gap between automated evaluation and actual BLV needs. First, we conduct an in-depth user study with BLV participants to understand and quantify their navigational preferences, curating VL-GUIDEDATA, a large-scale BLV user-simulated preference dataset containing image-request-response-score pairs. We then leverage the dataset to develop an accessibility-aware evaluator, VL-GUIDE-S, which outperforms existing (L)VLM judges in both human alignment and inference efficiency. Notably, its effectiveness extends beyond a single domain, demonstrating strong performance across multiple fine-grained, BLV-critical dimensions. We hope our work lays as a foundation for automatic AI judges that advance safe, barrier-free navigation for BLV users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。