arXiv:2605.19227cs.CRcs.AI2026-05

统一生成模型存在跨模态后门漏洞,一句普通词可触发虚假图文内容生成。

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models

论文配图:Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
图 1 · 摘自论文原文
  • 提出首个针对统一自回归模型的逐令牌后门攻击方法
  • 有模型访问时55%生成含品牌推广,无访问时数据投毒成功率63.1%
  • 用普通词汇作触发器,可同时操控图文输出,增强伪造内容可信度

统一自回归模型(UAMs)是将文本与图像标记在单一自回归过程中生成的Transformer模型。共享参数和多模态词汇表简化了训练流程并支持灵活的多模态生成,但可能引入新安全风险。我们首次揭示此类统一架构可引发跨模态后门攻击,即触发器可在多种输出模态间传播恶意影响。本文提出逐令牌后门攻击(ToBAC),探索基于数据与模型的污染策略。结果表明,无害字符甚至常见词汇可被转化为触发器,诱导自回归图像生成中出现有害行为。ToBAC能联合操控视觉输出与伴随文本,提升伪造内容的可信度。在有模型访问条件下,对统一液态模型(Liquid)的攻击中,仅需微小词语(如“cool”)即可在55%生成中引发对齐模态的品牌推广或意识形态影响;在无模型访问情况下,通过数据投毒亦可实现对JanusPro平均63.1%的成功率攻击。

原文摘要 · Abstract (English)

Unified autoregressive models (UAMs) are transformer models that generate text as well as image tokens within a single autoregressive pass. Shared parameters and a multimodal vocabulary simplify the training pipeline and facilitate flexible multimodal generation, yet might introduce new vulnerabilities. In particular, we are the first to show that this unified architecture enables multimodal backdoor attacks, where a trigger can propagate malicious effects across multiple output modalities. Specifically, we present the Token by Token Backdoor Attack (ToBAC), the first backdoor attack targeting UAMs, exploring both data-based and model-based poisoning strategies. We demonstrate that innocuous characters or even common words can be transformed into triggers that elicit harmful behavior in autoregressive image generation. ToBAC can jointly manipulate visual outputs and accompanying text, increasing the perceived authenticity of fabricated content. With model access, ToBAC enables attacks on the unified Liquid model in which a subtle word (e.g., ``cool'') induces modality-aligned brand promotion or ideological influence in 55% of generations. Without model access, ToBAC can be induced through data poisoning, achieving an average success rate of 63.1% against JanusPro.

后门攻击多模态生成模型安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。