用大模型模拟低视力者看图感知,发现需结合视觉信息和答题样例才准
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
- 构建40人问卷数据集,用大模型生成每位用户的模拟感知代理
- 仅给视觉能力或样例时模型准确率仅0.59,两者结合提升至0.70
- 一个含多选与开放题的样例最有效,再多也无明显提升
视觉语言模型(VLMs)在模拟人类行为推理方面取得进展,但其在无障碍领域的应用尚未被研究。本文评估了大模型(GPT-4o)在模拟低视力人群图像感知方面的表现。通过40名低视力参与者完成的问卷调查,收集其简要与详细视觉信息,以及对最多25张图像的开放式和多项选择式感知与识别回答,构建用于VLM的提示。我们生成每位参与者的模拟代理,测试不同输入条件下模型输出与原始回答的一致性。结果显示:在仅有最小提示时,模型常超出设定视力能力,一致性仅为0.59;仅提供视觉信息或样例时,一致性仍为0.59;而两者结合可显著提升至0.70(p < 0.0001)。值得注意的是,单个包含开放题与多选题的样例相比单一类型有显著提升(p < 0.0001),额外样例则无明显增益(p > 0.05)。
原文摘要 · Abstract (English)
Advances in vision language models (VLMs) have enabled the simulation of general human behavior through their reasoning and problem solving capabilities. However, prior research has not investigated such simulation capabilities in the accessibility domain. In this paper, we evaluate the extent to which VLMs can simulate the vision perception of low vision individuals when interpreting images. We first compile a benchmark dataset through a survey study with 40 low vision participants, collecting their brief and detailed vision information and both open-ended and multiple-choice image perception and recognition responses to up to 25 images. Using these responses, we construct prompts for VLMs (GPT-4o) to create simulated agents of each participant, varying the included information on vision information and example image responses. We evaluate the agreement between VLM-generated responses and participants' original answers. Our results indicate that VLMs tend to infer beyond the specified vision ability when given minimal prompts, resulting in low agreement (0.59). The agreement between the agent' and participants' responses remains low when only either the vision information (0.59) or example image responses (0.59) are provided, whereas a combination of both significantly increase the agreement (0.70, p < 0.0001). Notably, a single example combining both open-ended and multiple-choice responses, offers significant performance improvements over either alone (p < 0.0001), while additional examples provided minimal benefits (p > 0.05).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。