研究视觉语言模型在图像与文本特征变化下的表现差异。
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features
- 构建多维度实验框架,分析图像和提示对模型的影响
- 微小修改即引发模型回答显著变化,尤其在提示具体时
- 适合关注模型可靠性与偏差的开发者和研究者
近期研究表明,视觉语言模型(VLMs)在回答图像属性问题时依赖训练中习得的固有偏见。当被问及需要聚焦图像特定区域的高细节问题时,这种偏见会被放大。例如,要求统计修改版美国国旗上的星星数量(如超过50颗),模型常忽略视觉证据而答错。本文基于此,提出多维度实验框架,系统分析输入数据特征(图像与提示)如何影响性能。使用开源VLMs,研究注意力值随图像尺寸、物体数量、背景色、提示具体性等参数的变化。结果表明,即使细微的图像特征或提示具体性调整,也会导致模型回答方式与整体性能产生显著差异。
原文摘要 · Abstract (English)
Recent research on Vision Language Models (VLMs) suggests that they rely on inherent biases learned during training to respond to questions about visual properties of an image. These biases are exacerbated when VLMs are asked highly specific questions that require focusing on specific areas of the image. For example, a VLM tasked with counting stars on a modified American flag (e.g., with more than 50 stars) will often disregard the visual evidence and fail to answer accurately. We build upon this research and develop a multi-dimensional examination framework to systematically determine which characteristics of the input data, including both the image and the accompanying prompt, lead to such differences in performance. Using open-source VLMs, we further examine how attention values fluctuate with varying input parameters (e.g., image size, number of objects in the image, background color, prompt specificity). This research aims to learn how the behavior of vision language models changes and to explore methods for characterizing such changes. Our results suggest, among other things, that even minor modifications in image characteristics and prompt specificity can lead to large changes in how a VLM formulates its answer and, subsequently, its overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。