arXiv:2510.03612cs.AIcs.CR2025-10

用视觉和文本联合干扰,悄悄操纵网页代理的推荐结果。

Cross-Modal Content Optimization for Steering Web Agent Preferences

  • 通过同时修改图片和描述文本,实现隐蔽的偏好操控。
  • 在多个主流模型上均显著优于基线,且检测率低70%。
  • 适合关注AI安全与对抗攻击的研究者或从业者。

基于视觉-语言模型(VLM)的网页代理正广泛应用于内容推荐、商品排序等高风险任务,结合多模态感知与偏好推理能力。然而,现有研究揭示这些代理易受攻击:攻击者可通过对抗性弹窗、图像扰动或内容篡改来扭曲选择结果。以往工作通常假设具备强白盒访问权限,或仅限单模态扰动,或采用不现实的设置。本文首次证明,在真实攻击者条件下,联合利用视觉与文本通道可带来更强大的偏好操控。我们提出跨模态偏好引导(CPS)方法,对项目图像与自然语言描述进行不可察觉的联合优化,利用可迁移的图像扰动与基于强化学习的人类反馈(RLHF)引发的语言偏见,引导代理决策。与依赖梯度访问、网页控制或代理记忆的先前方法不同,本工作采用真实黑盒威胁场景:非特权攻击者仅可修改自身条目的图像和文本元数据,无任何模型内部信息。我们在GPT-4.1、Qwen-2.5VL和Pixtral-Large等先进专有及开源VLM驱动的代理上评估了电影选择与电商任务。结果显示,CPS在所有模型上均显著优于主流基线方法,且检测率降低70%,兼具高效性与隐蔽性。这些发现凸显了在代理系统日益重要的社会背景下,亟需构建鲁棒防御机制。

原文摘要 · Abstract (English)

Vision-language model (VLM)-based web agents increasingly power high-stakes selection tasks like content recommendation or product ranking by combining multimodal perception with preference reasoning. Recent studies reveal that these agents are vulnerable against attackers who can bias selection outcomes through preference manipulations using adversarial pop-ups, image perturbations, or content tweaks. Existing work, however, either assumes strong white-box access, with limited single-modal perturbations, or uses impractical settings. In this paper, we demonstrate, for the first time, that joint exploitation of visual and textual channels yields significantly more powerful preference manipulations under realistic attacker capabilities. We introduce Cross-Modal Preference Steering (CPS) that jointly optimizes imperceptible modifications to an item's visual and natural language descriptions, exploiting CLIP-transferable image perturbations and RLHF-induced linguistic biases to steer agent decisions. In contrast to prior studies that assume gradient access, or control over webpages, or agent memory, we adopt a realistic black-box threat setup: a non-privileged adversary can edit only their own listing's images and textual metadata, with no insight into the agent's model internals. We evaluate CPS on agents powered by state-of-the-art proprietary and open source VLMs including GPT-4.1, Qwen-2.5VL and Pixtral-Large on both movie selection and e-commerce tasks. Our results show that CPS is significantly more effective than leading baseline methods. For instance, our results show that CPS consistently outperforms baselines across all models while maintaining 70% lower detection rates, demonstrating both effectiveness and stealth. These findings highlight an urgent need for robust defenses as agentic systems play an increasingly consequential role in society.

AI安全对抗攻击跨模态推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。