arXiv:2502.14883cs.CVcs.AI2025-02被引 1

研究视障者对大模型生成场景描述的偏好,提升无障碍体验。

How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions

  • 通过用户调研对比六类大模型描述,发现效果差异大。
  • 参与者普遍认为描述能降低恐惧感,但简洁性评分参差不齐。
  • 提出新评估指标,需结合用户反馈优化模型输出。

视盲或低视力(BLV)人群在复杂环境中导航面临严重风险。大型视觉语言模型(LVLM)在生成场景描述方面展现出潜力,但其对BLV用户的实际效用仍待探索。为填补这一空白,我们对8名BLV参与者进行了用户研究,系统评估了六种类型LVLM描述的偏好。尽管这些描述有助于减轻恐惧感并提升可操作性,但用户对描述充分性和简洁性的评分存在显著差异。此外,尽管GPT-4o具备强大的描述优化能力,却未获得参与者一致青睐。基于用户研究结果,我们构建了训练数据,用于开发能够有效捕捉BLV偏好的自动评估指标。研究强调,亟需以BLV为中心的评估体系与人机协同反馈机制,以推动LVLM描述质量在无障碍领域的持续进步。

原文摘要 · Abstract (English)

For individuals with blindness or low vision (BLV), navigating complex environments can pose serious risks. Large Vision-Language Models (LVLMs) show promise for generating scene descriptions, but their effectiveness for BLV users remains underexplored. To address this gap, we conducted a user study with eight BLV participants to systematically evaluate preferences for six types of LVLM descriptions. While they helped to reduce fear and improve actionability, user ratings showed wide variation in sufficiency and conciseness. Furthermore, GPT-4o--despite its strong potential to refine descriptions--was not consistently preferred by participants. We use the insights obtained from the user study to build training data for building our new automatic evaluation metric that can capture BLV preferences effectively. Our findings underscore the urgent need for BLV-centered evaluation metrics and human-in-the-loop feedback to advance LVLM description quality for accessibility.

无障碍视觉语言模型用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。