颜色影响模型判断,绿色让正面词更正向,引发视觉偏见。
Seeing Red, Thinking Bad: Color Bias in Vision Language Models

- 用隐蔽视觉提示改变文字颜色对比,保持语义不变
- 绿色使正面词更正向,负词易被忽略,准确率下降15%以上
- 适合关注AI视觉偏见、模型可解释性的研究者
视觉语言模型(VLMs)在招聘辅助、推荐系统等工业决策中广泛应用,促使我们深入分析其对图文信息的处理机制。本文研究文本以图像形式呈现时,视觉风格(如颜色、对比度)如何影响模型理解。为此,我们提出隐蔽视觉提示(Stealth Visual Prompts),在不改变语义的前提下,微妙调整文本的视觉样式。通过系统性控制文字视觉特征,我们发现将正面词汇渲染为绿色会显著使情感判断偏向正面;同时,降低文字与背景对比度会增加模型对视觉显著性线索的依赖,导致视觉问答(VQA)错误率上升15%以上。进一步分析表明,这些偏差与视觉编码器的潜在表示变化密切相关。结果说明,文本视觉样式可能引导模型理解偏离人类语义认知,存在潜在风险。
原文摘要 · Abstract (English)
Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how VLMs interpret text rendered as an image, and investigate the influence of visual styling biases. To this end, we introduce Stealth Visual Prompts, which subtly change visual styling of text, such as color and contrast, while preserving semantic content. Using these prompts, we systematically control the visual styling of words in text and measure their impact on the analysis performed by VLMs. We further analyze how such visual perturbations affect the latent representations of the vision encoder. From our experiments, we observed that coloring positive words in green consistently shifts sentiment predictions toward a positive direction. As a result, VLMs often fail to properly account for negative words present in the text. Our analysis suggests that this behavior is correlated with changes in the latent representations of the vision encoder induced by color variations. In addition, we show that reducing text--background contrast increases reliance on visually salient cues and leads to more incorrect Visual Question Answering (VQA) outputs. These results suggest that the visual styling of rendered text can guide VLMs' interpretation in ways that diverge from human semantic understanding. Project page: https://github.com/KohsukeIde/color-bias-vlm
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。