图像提示能改变视觉语言模型的合作行为,影响其决策。
The Effects of Visual Priming on Cooperative Behavior in Vision-Language Models

- 用图像和颜色线索测试模型在囚徒困境中的合作倾向。
- 不同模型对视觉提示反应差异大,部分可被提示或推理缓解。
- 适合关注AI安全与视觉输入影响的研究者阅读。
随着视觉语言模型(VLMs)越来越多地融入决策系统,理解视觉输入如何影响其行为至关重要。本文以迭代囚徒困境(IPD)为测试场景,研究图像提示对VLM合作行为的影响。通过展示体现友善/助人或攻击/自私行为概念的图像,以及彩色奖励矩阵,考察其对模型决策模式的作用。实验在多个领先VLM上进行,并探索了提示修改、思维链(CoT)推理和视觉令牌削减等缓解策略。结果表明,图像内容与颜色线索均能影响VLM行为,不同模型对此的敏感度及缓解效果各异。这些发现强调了在视觉丰富且安全关键环境中部署VLM时,建立鲁棒评估框架的重要性,同时也揭示了模型架构与训练差异可能导致截然不同的行为响应,值得进一步研究。
原文摘要 · Abstract (English)
As Vision-Language Models (VLMs) become increasingly integrated into decision-making systems, it is essential to understand how visual inputs influence their behavior. This paper investigates the effects of visual priming on VLMs' cooperative behavior using the Iterated Prisoner's Dilemma (IPD) as a test scenario. We examine whether exposure to images depicting behavioral concepts (kindness/helpfulness vs. aggressiveness/selfishness) and color-coded reward matrices alters VLM decision patterns. Experiments were conducted across multiple state-of-the-art VLMs. We further explore mitigation strategies including prompt modifications, Chain of Thought (CoT) reasoning, and visual token reduction. Results show that VLM behavior can be influenced by both image content and color cues, with varying susceptibility and mitigation effectiveness across models. These findings not only underscore the importance of robust evaluation frameworks for VLM deployment in visually rich and safety-critical environments, but also highlight how architectural and training differences among models may lead to distinct behavioral responses-an area worthy of further investigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。