arXiv:2607.02819cs.CRcs.AI2026-07

攻击者仅篡改10%视觉令牌,即让大模型推理准确率下降88%

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models

论文配图:Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models
图 1 · 摘自论文原文
  • 黑盒中间人攻击,篡改传输中的视觉令牌
  • 仅改10%令牌,最高使准确率下降88.31%
  • 揭示云边协同视觉语言模型的致命漏洞

云边协同的大规模视觉语言模型(LVLM)通过在边缘设备与云端服务器之间分割计算实现高效部署。在此过程中,中间视觉令牌需经通信链路从边缘传至云端,形成新的攻击面。本文研究在黑盒中间人场景下的视觉令牌操纵攻击(VTM-Attack),攻击者在预算约束下拦截并篡改部分传输的视觉令牌。提出四种简单攻击策略及一种基于优化的令牌选择方法。在6个主流LVLM(3B-72B)和4个基准测试上的实验表明,仅操纵10%的视觉令牌即可导致准确率最高下降88.31%。结果揭示了云边协同LVLM推理中的关键安全漏洞。

原文摘要 · Abstract (English)

Cloud-edge Large Vision-Language Model (LVLM) inference enables efficient deployment by splitting computation between edge devices and cloud servers. In this process, intermediate vision tokens are transmitted from the edge to the cloud over a communication link, thereby exposing a new attack surface. We study vision token manipulation attack (VTM-Attack) under a black-box man-in-the-middle setting, where an adversary intercepts and manipulates a subset of transmitted vision tokens under a budget constraint. We propose four naïve attack strategies and an optimization-based token selection method. Experiments on 6 state-of-the-art LVLMs (3B-72B) across 4 benchmarks show that manipulating only 10\% of vision tokens can reduce accuracy by up to 88.31\%. These results reveal a critical vulnerability in cloud-edge LVLM inference.

视觉语言模型云边协同安全攻击令牌操纵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。