通过伪造导向增强,让检测器聚焦生成视频的底层痕迹。
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation
- 用小波分解替换频段,强化模型对生成痕迹的感知。
- 仅用单一模型数据训练,跨多个生成模型测试准确率显著提升。
- 适合需要高泛化能力的AI视频真伪检测场景。
合成视频生成技术快速发展,最新模型可生成与真实视频几乎无法区分的高清视频。尽管已有多种视频取证检测方法,但普遍泛化能力差,难以应用于实际场景。本文的核心洞察是:优秀的检测器应关注生成架构引入的内在低层痕迹,而非特定模型的高层语义缺陷。首先,研究不同生成架构,识别出无偏、鲁棒且跨模型共享的判别性特征;其次,提出一种基于小波分解的伪造导向数据增强策略,通过替换特定频段来引导模型捕捉更相关的取证线索。该训练范式无需复杂算法或大规模多生成器数据集,即可显著提升检测器的泛化能力。在实验中,仅使用单个生成模型数据训练,测试时覆盖多种其他生成模型,结果显著优于现有方法,并在NOVA和FLUX等最新模型上仍表现优异。
原文摘要 · Abstract (English)
Synthetic video generation is progressing very rapidly. The latest models can produce very realistic high-resolution videos that are virtually indistinguishable from real ones. Although several video forensic detectors have been recently proposed, they often exhibit poor generalization, which limits their applicability in a real-world scenario. Our key insight to overcome this issue is to guide the detector towards *seeing what really matters*. In fact, a well-designed forensic classifier should focus on identifying intrinsic low-level artifacts introduced by a generative architecture rather than relying on high-level semantic flaws that characterize a specific model. In this work, first, we study different generative architectures, searching and identifying discriminative features that are unbiased, robust to impairments, and shared across models. Then, we introduce a novel forensic-oriented data augmentation strategy based on the wavelet decomposition and replace specific frequency-related bands to drive the model to exploit more relevant forensic cues. Our novel training paradigm improves the generalizability of AI-generated video detectors, without the need for complex algorithms and large datasets that include multiple synthetic generators. To evaluate our approach, we train the detector using data from a single generative model and test it against videos produced by a wide range of other models. Despite its simplicity, our method achieves a significant accuracy improvement over state-of-the-art detectors and obtains excellent results even on very recent generative models, such as NOVA and FLUX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。