让视觉语言模型学会判断图像是否够用,并给出调整建议。
Right this way: Can VLMs Guide Us to See More to Answer Questions?
- 提出新任务:让VLM判断图像信息是否不足并指导如何调整
- 在合成数据上微调后,主流VLM性能显著提升
- 对视障人士获取有效图像有实际帮助
在问答场景中,人类能评估信息是否充足,必要时主动获取更多信息,而非强行作答。而当前视觉语言模型(VLMs)通常直接生成一次性回答,不评估信息充分性。为探究这一差距,我们识别出视觉问答(VQA)中的关键挑战:当视觉信息不足以回答问题时,VLM能否指示应如何调整图像?这一能力对视障人士正确拍摄图像尤为关键。为此,我们构建了一个人工标注的数据集作为该任务的基准。此外,我们提出一种自动化框架,通过模拟‘何处需知’场景生成合成训练数据。实验证明,在该合成数据上微调后,主流VLM性能显著提升。本研究展示了缩小VLM在信息评估与获取之间差距的潜力,使其表现更接近人类。
原文摘要 · Abstract (English)
In question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answer. In contrast, Vision Language Models (VLMs) typically generate direct, one-shot responses without evaluating the sufficiency of the information. To investigate this gap, we identify a critical and challenging task in the Visual Question Answering (VQA) scenario: can VLMs indicate how to adjust an image when the visual information is insufficient to answer a question? This capability is especially valuable for assisting visually impaired individuals who often need guidance to capture images correctly. To evaluate this capability of current VLMs, we introduce a human-labeled dataset as a benchmark for this task. Additionally, we present an automated framework that generates synthetic training data by simulating ``where to know'' scenarios. Our empirical results show significant performance improvements in mainstream VLMs when fine-tuned with this synthetic data. This study demonstrates the potential to narrow the gap between information assessment and acquisition in VLMs, bringing their performance closer to humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。