arXiv:2603.10504cs.CRcs.AI2026-03

用聊天机器人轻松生成骗过检测的高质伪造图像

Naïve Exposure of Generative AI Capabilities Undermines Deepfake Detection

  • 利用商用AI聊天机器人无害提示实现图像精修
  • 精修后图像既避检又保真且画质显著提升
  • 普通用户也能用现成工具绕过主流检测系统

生成式AI系统通过面向用户的聊天机器人界面暴露了强大的推理与图像优化能力。本文表明,此类能力的随意暴露会从根本上削弱现有深度伪造检测技术。我们研究了一种真实且已部署的攻击场景:攻击者仅使用合规提示和商业生成式AI系统,不涉及新型篡改技术。结果表明,当前最先进的深度伪造检测方法在语义保持的图像精修下失效。具体而言,生成式AI系统会显式表达真实性标准,并通过不受限的推理过程无意中将其外化为优化目标。由此生成的图像同时满足:逃避检测、经商用人脸识别API验证身份一致、且感知质量显著提升。重要的是,我们发现广泛可用的商业聊天服务比开源模型带来更大安全风险,因其具备更高的真实感、语义可控性及低门槛交互,使非专业用户也能有效绕过检测。研究揭示了当前检测框架所假设的威胁模型与真实世界生成式AI能力之间存在结构性错配。

原文摘要 · Abstract (English)

Generative AI systems increasingly expose powerful reasoning and image refinement capabilities through user-facing chatbot interfaces. In this work, we show that the naïve exposure of such capabilities fundamentally undermines modern deepfake detectors. Rather than proposing a new image manipulation technique, we study a realistic and already-deployed usage scenario in which an adversary uses only benign, policy-compliant prompts and commercial generative AI systems. We demonstrate that state-of-the-art deepfake detection methods fail under semantic-preserving image refinement. Specifically, we show that generative AI systems articulate explicit authenticity criteria and inadvertently externalize them through unrestricted reasoning, enabling their direct reuse as refinement objectives. As a result, refined images simultaneously evade detection, preserve identity as verified by commercial face recognition APIs, and exhibit substantially higher perceptual quality. Importantly, we find that widely accessible commercial chatbot services pose a significantly greater security risk than open-source models, as their superior realism, semantic controllability, and low-barrier interfaces enable effective evasion by non-expert users. Our findings reveal a structural mismatch between the threat models assumed by current detection frameworks and the actual capabilities of real-world generative AI. While detection baselines are largely shaped by prior benchmarks, deployed systems expose unrestricted authenticity reasoning and refinement despite stringent safety controls in other domains.

深度伪造AI安全检测对抗生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。