arXiv:2501.19143cs.AIcs.CR2025-01

用类模仿游戏和思维链机制,同时防御生成模型的两类对抗幻觉。

Imitation Game for Adversarial Disillusion with Chain-of-Thought Reasoning in Generative AI

论文配图:Imitation Game for Adversarial Disillusion with Chain-of-Thought Reasoning in Generative AI
图 1 · 摘自论文原文
  • 设计多模态生成代理,通过思维链理解并重构样本语义本质。
  • 在白盒与黑盒攻击下,成功消除推导型与归纳型对抗幻觉。
  • 适合关注生成模型安全与鲁棒性验证的研究者参考。

作为人工智能核心的机器感知正面临对抗幻觉的根本威胁。这类攻击主要分为两类:推导型幻觉,即基于目标模型的通用决策逻辑构造特定刺激以干扰其判断;归纳型幻觉,即通过特定刺激在学习阶段植入后门,触发时导致异常行为。二者共同揭示了防御体系的复杂性,亟需统一应对策略。本文提出一种基于模仿游戏的祛魅范式,核心为一个由思维链驱动的多模态生成代理,能观察、内化并重建样本的语义本质,不再局限于还原原始样本。作为概念验证,我们在多模态生成对话代理上进行实验,评估该方法在多种攻击场景下的表现。结果表明,所提框架在各类白盒与黑盒攻击中均能有效中和推导型与归纳型对抗幻觉。

原文摘要 · Abstract (English)

As the cornerstone of artificial intelligence, machine perception confronts a fundamental threat posed by adversarial illusions. These adversarial attacks manifest in two primary forms: deductive illusion, where specific stimuli are crafted based on the victim model's general decision logic, and inductive illusion, where the victim model's general decision logic is shaped by specific stimuli. The former exploits the model's decision boundaries to create a stimulus that, when applied, interferes with its decision-making process. The latter reinforces a conditioned reflex in the model, embedding a backdoor during its learning phase that, when triggered by a stimulus, causes aberrant behaviours. The multifaceted nature of adversarial illusions calls for a unified defence framework, addressing vulnerabilities across various forms of attack. In this study, we propose a disillusion paradigm based on the concept of an imitation game. At the heart of the imitation game lies a multimodal generative agent, steered by chain-of-thought reasoning, which observes, internalises and reconstructs the semantic essence of a sample, liberated from the classic pursuit of reversing the sample to its original state. As a proof of concept, we conduct experimental simulations using a multimodal generative dialogue agent and evaluates the methodology under a variety of attack scenarios. Experimental results demonstrate that the proposed framework consistently neutralises both deductive and inductive adversarial illusions across diverse white-box and black-box attack scenarios.

对抗攻击生成模型思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。