通过可控实验揭示视觉语言模型的偏好机制。
Visual Persuasion: What Influences Decisions of Vision-Language Models?
- 用图像编辑+对比选择,反推模型视觉偏好
- 优化后的图像可显著提升被选概率(↑37%)
- 适合研究AI决策安全与可解释性的人群
网络上充斥着为人类设计的图像,如今越来越多由使用视觉语言模型(VLMs)的智能体进行解读。这些智能体在大规模场景下做出视觉决策,如点击、推荐或购买。然而,我们对它们的视觉偏好结构知之甚少。本文提出一种研究框架,将VLM置于受控的图像选择任务中,并系统性地扰动输入。核心思路是将智能体的决策函数视为潜在的视觉效用,通过揭示偏好(即在系统化修改图像后的选择行为)来推断。从常见图像(如商品图)出发,我们提出视觉提示优化方法,借鉴文本优化技术,利用图像生成模型(如调整构图、光照或背景)迭代生成视觉上合理的修改。随后评估哪些修改能提升被选概率。在前沿VLM上的大规模实验表明,优化后的修改在一对一比较中显著改变选择概率。我们还开发了自动可解释性流程,识别出驱动选择的一致视觉主题。该方法提供了一种高效、实用的途径,以发现视觉漏洞和潜在安全风险,这些风险原本只能在真实环境中被动暴露,从而支持对图像基AI代理的主动审计与治理。
原文摘要 · Abstract (English)
The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (VLMs). These agents make visual decisions at scale, deciding what to click, recommend, or buy. Yet, we know little about the structure of their visual preferences. We introduce a framework for studying this by placing VLMs in controlled image-based choice tasks and systematically perturbing their inputs. Our key idea is to treat the agent's decision function as a latent visual utility that can be inferred through revealed preference: choices between systematically edited images. Starting from common images, such as product photos, we propose methods for visual prompt optimization, adapting text optimization methods to iteratively propose and apply visually plausible modifications using an image generation model (such as in composition, lighting, or background). We then evaluate which edits increase selection probability. Through large-scale experiments on frontier VLMs, we demonstrate that optimized edits significantly shift choice probabilities in head-to-head comparisons. We develop an automatic interpretability pipeline to explain these preferences, identifying consistent visual themes that drive selection. We argue that this approach offers a practical and efficient way to surface visual vulnerabilities, safety concerns that might otherwise be discovered implicitly in the wild, supporting more proactive auditing and governance of image-based AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。