arXiv:2607.15246cs.CV2026-07

用多智能体框架提升伪造图像逃逸检测的跨模型攻击能力

ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors

论文配图:ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
图 1 · 摘自论文原文
  • 通过视觉语言模型提供空间语义先验,智能调度多种扰动策略
  • 在AADD-2025上对齐高质量与低质量图像均实现超前攻击成功率
  • 适合研究深度伪造防御漏洞及对抗攻击的从业者参考

深度伪造检测器在黑盒对抗迁移下可靠性常下降,因其依赖脆弱且架构相关的取证线索。现有迁移攻击缺乏语义感知,在严格无查询约束下效果不佳,尤其当扰动从卷积代理迁移到基于变压器的目标时。本文提出ARMOR++,一种鲁棒的多智能体框架,用于高迁移性深度伪造逃逸攻击。该框架利用Qwen2.5-VL视觉语言模型(VLM)提供空间语义先验,Qwen3大语言模型(LLM)负责原始基元选择、自适应超参数重参数化及熵正则化扰动混合。通过整合五种互补基元——密集优化、显著性方法、空间变换、频域扰动和块结构修改——ARMOR++有效针对异构归纳偏置。在AADD-2025基准上的严格评估显示,其显著优于现有代理与非代理基线,覆盖低/高质量图像场景。统计分析证实其盲目标攻击成功率(ASR)远超当前最优代理基线,且在强防御配置下仍具优势。这些发现揭示了当前深度伪造检测部署中的显著可靠性缺口,并验证了代理协同在识别潜在漏洞方面的有效性。

原文摘要 · Abstract (English)

The reliability of deepfake detectors frequently degrades under black-box adversarial transfer, as these models often rely on fragile, architecture-dependent forensic cues. Existing transfer attacks often lack semantic awareness and struggle to maintain effectiveness under strict no-query constraints, particularly when perturbations are transferred from convolutional surrogates to transformer-based targets. To address these limitations, this paper introduces ARMOR++, a robust multi-agent framework designed for high-transferability deepfake evasion. The framework leverages the Qwen2.5-VL Vision-Language Model (VLM) to supply spatial semantic priors, while the Qwen3 Large Language Model (LLM) orchestrates primitive selection, adaptive hyperparameter reparameterization, and entropy-regularized perturbation mixing. By integrating five complementary primitives, spanning dense optimization, saliency-based methods, spatial transformations, frequency-domain perturbations, and block-structured modifications, ARMOR++ effectively targets heterogeneous inductive biases. Rigorous evaluation on the AADD-2025 benchmark demonstrates that ARMOR++ significantly outperforms existing agentic and non-agentic baselines across both low- and high-quality image regimes. Statistical analysis confirms a substantial gain in blind-target Attack Success Rate (ASR) over the state-of-the-art agentic baseline, with further performance advantages evidenced against non-agentic benchmarks and under robust defensive configurations. These findings highlight a significant residual reliability gap in current deepfake detector deployments and demonstrate the efficacy of agentic orchestration in identifying latent vulnerabilities.

对抗攻击深度伪造多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。