测试视觉语言模型在多模态说服中的脆弱性,发现图像比文字更易绕过安全防御。
Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion
- 构建多模态说服框架MMPersuade,模拟视觉与心理策略的联合影响
- 图像输入使模型在对抗场景中更易被说服,突破文本安全机制
- 不同领域和格式下模型脆弱性差异大,高能力模型反而在对抗中更易受骗
随着自主代理间交互增多,它们不断尝试相互影响。以往研究聚焦于纯文本环境下的代理间(A2A)说服,但视觉语言模型(VLMs)的兴起带来了新挑战:多模态内容蕴含更丰富信息,同时嵌入微妙且难以察觉的说服线索。为此,我们提出MMPersuade——一个统一的框架与数据集,用于研究A2A多模态说服。我们建模了利用图像与心理策略的说服者代理,与接受说服的VLM之间的互动。基准涵盖商业、主观行为及对抗性情境,通过函数调用评估说服效果,捕捉行为变化而非仅言语响应。在六种VLM上的实验揭示三个发现:(1)多模态输入持续优于纯文本说服,原始视觉信号在对抗场景中显著提升受骗率,因可绕过文本激活的安全机制;(2)受说服程度高度依赖领域与格式,在商业场景中真实感与社区风格格式更具影响力,而在对抗场景中则由不同格式主导;(3)心理策略有效性随上下文与模型架构而变,更强大的模型能抵抗良性说服,但在对抗性多模态输入下反而更易受影响。该框架为构建更鲁棒、对齐的多代理VLM奠定了基础。
原文摘要 · Abstract (English)
As autonomous agents increasingly interact, they inevitably attempt to influence one another. While prior work in text-only settings has explored the dynamics of Agent-to-Agent (A2A) persuasion, the rise of Vision-Language Models (VLMs) introduces a more complex challenge: multimodal content conveys richer information while integrating subtle, hard-to-detect persuasive cues. To study this vulnerability, we present MMPersuade, a unified framework and dataset for A2A multimodal persuasion. We model interactions between a persuader agent, which leverages images and psychological strategies, and a persuadee VLM. Our benchmark spans commercial, subjective and behavioral, and adversarial contexts, and evaluates persuasion via function-calling that capture behavioral shifts beyond verbal responses. Experiments on six VLMs reveal three findings: (1) multimodal inputs consistently outperform text-only persuasion, with raw visual signals uniquely increasing susceptibility in adversarial settings by bypassing text-activated safety defenses; (2) persuadee vulnerability is highly domain- and format-dependent, with realistic and community-style formats driving susceptibility in commercial settings while different formats dominate in adversarial ones; and (3) psychological strategy efficacy varies with context and model architecture, as more capable models resist benign persuasion yet become more susceptible under adversarial multimodal inputs. Our framework provides a foundation for building more robust and aligned VLMs in multi-agent environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。