arXiv:2601.12263cs.CLcs.AI2026-01被引 5

攻击者可微调图片和文字,让视觉语言模型错排序商品。

Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers

  • 通过联合扰动图像和文本,操控模型内部跨模态关联。
  • 攻击效果远超单模态方法,且在商用大模型上有效。
  • 揭示了多模态模型排名机制的脆弱性,适合安全研究者参考。

视觉语言模型(VLM)将视觉与文本知识融合为统一表征,广泛应用于现代检索与推荐系统。然而,其在多模态项目排序中对跨模态知识的利用是否可靠,以及是否存在被攻破的可能仍不明确。本文揭示了一种基础性漏洞:通过多模态生成引擎优化(MGEO),攻击者可同时生成难以察觉的图像扰动与流畅的文本后缀,利用模型内部跨模态知识耦合机制操纵排序结果。采用交替优化策略,MGEO针对VLM内部视觉与语言表示的深层交互,实现的排名操纵效果显著优于单模态攻击和基于强商业模型的启发式基线。结果表明,仅提升表面内容质量不足以促进排名,必须直接对齐模型内部知识利用机制。该发现引发对多模态基础模型知识锚定忠实性与鲁棒性的深刻质疑,并推动未来防御机制的研究。代码已开源:https://github.com/glad-lab/MGEO

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) integrate visual and textual knowledge into unified representations that increasingly underpin modern retrieval and recommendation systems. However, it remains unclear how reliably these models utilize their cross-modal knowledge when ranking multimodal items, and whether their knowledge grounding can be subverted. In this paper, we expose a fundamental vulnerability in how VLMs apply multimodal knowledge for product ranking: through Multimodal Generative Engine Optimization (MGEO), we show that an adversary can manipulate a VLM's ranking decisions by jointly crafting imperceptible image perturbations and fluent textual suffixes that exploit the model's internal cross-modal knowledge coupling. Using an alternating optimization strategy, MGEO targets the deep interactions between visual and linguistic representations within the VLM, achieving rank manipulations that substantially exceed those of unimodal attacks and heuristic baselines powered by strong commercial models. Our findings reveal that surface-level content quality is insufficient for rank promotion; instead, direct alignment with the model's internal knowledge utilization mechanism is required. These results raise important questions on the faithfulness and robustness of knowledge grounding in multimodal foundation models, and motivate future work on defense mechanisms for multimodal retrieval systems. Code is available at: https://github.com/glad-lab/MGEO

多模态对抗攻击视觉语言模型排序安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。