arXiv:2505.13302cs.CL2025-05AAAI

图像让假新闻更易被模型转发,且不同模型和人格设定影响程度不同。

Images Amplify Misinformation Sharing in Vision-Language Models

  • 用诱导式提示绕过模型拒绝,测试其转发决策行为。
  • 有图时假新闻转发率上升14.5%,真新闻升5.3%。
  • 性格设定如黑暗三联征会加剧假新闻传播,适合关注AI风险的研究者。

随着语言与视觉语言模型(VLMs)成为信息获取与在线互动的核心,其放大虚假信息的潜在风险日益引发关注。人类研究表明,图像能提升信息的可信度与传播性,由此提出疑问:VLMs 是否也存在类似脆弱性?我们首次研究了图像如何影响 VLMs 的新闻重分享倾向,以及该效应在不同模型家族间的差异,及人格设定与内容属性的调节作用。我们设计了一种受“越狱”启发的提示策略,突破 VLMs 对争议新闻的默认拒绝,使其能在多样话题与人格特征(包括反社会特质)下生成转发决策。我们在一个新构建的多模态数据集上评估了四个顶尖 VLMs,该数据集来自 PolitiFact 的经过核实的政治新闻,配以图像和真实性的标签。实验显示,图像存在使假新闻的转发率提升14.5%,真新闻提升5.3%。人格设定进一步调节该效应:黑暗三联征特质会加剧假新闻的转发,而亲共和党倾向的设定则降低对真实性的敏感度。在测试模型中,Claude-3-Haiku 表现出最强的抗视觉误导能力。结果表明,VLMs 复现了人类对图像的偏见,凸显多模态 AI 系统的新兴风险。研究呼吁建立考虑视觉影响与人格驱动变异性的评估框架与缓解策略,尤其在人工智能塑造公众话语与信息传播的社会技术场景中。

原文摘要 · Abstract (English)

As language and vision-language models (VLMs) become central to information access and online interaction, concerns grow about their potential to amplify misinformation. Human studies show that images boost the perceived credibility and shareability of information, raising the question of whether VLMs exhibit the same vulnerability. We present the first study examining how images influence VLMs' propensity to reshare news content, how this effect varies across model families, and how persona conditioning and content attributes modulate such behavior. We develop a jailbreaking-inspired prompting strategy that bypasses VLMs' default refusals to engage with controversial news, allowing them to generate resharing decisions across diverse topics and elicited traits, including antisocial ones. We evaluate four state-of-the-art VLMs on a novel multimodal dataset of fact-checked political news from PolitiFact, paired with images and ground-truth veracity labels. Our experiments show that image presence increases resharing rates by 14.5% for false news and 5.3% for true news. Persona conditioning further modulates this effect: Dark Triad traits amplify resharing of false news, whereas Republican-aligned profiles reduce sensitivity to veracity. Among the tested models, Claude-3-Haiku demonstrates the greatest robustness to visual misinformation. These findings reveal that VLMs replicate human-like biases in response to images, underscoring emerging risks for multimodal AI systems. They point to the need for evaluation frameworks and mitigation strategies that account for visual influence and persona-driven variability, particularly in sociotechnical settings where AI systems shape public discourse and information sharing.

视觉语言模型虚假信息人格影响传播风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。