通过放大噪声信号,精准识别文本生成视频的伪造痕迹。
Revealing Artifacts via Noise Amplification: A Novel Perspective for AI-Generated Video Detection

- 从比特平面出发,提取并放大视频中的细微噪声信号。
- 在GenVidBench和HardGVD上超越现有方法,准确率显著提升。
- 适合关注生成视频检测、对抗伪造内容的研究者使用。
随着视频生成模型的快速发展,区分人工智能生成视频与真实视频成为一项挑战。现有研究多聚焦于生成对抗网络生成样本的检测,而针对文本到视频模型生成的视频检测仍属空白。尽管先进文本到视频模型能生成逼真视觉内容,但在图像细节及视频内细节变化方面仍存在不足。受此启发,本文提出一种基于比特平面的新视角,可有效刻画图像或视频中的细节与噪声。为此,我们设计了一种简单而高效的噪声放大方法:首先基于比特平面提取噪声信号,再进行放大,最后输入判别网络进行视频伪造分类。该方法融合像素级强度增强、区域级空间放大与帧级时间聚合三个层面。为评估复杂场景下的检测性能,我们还构建了名为HardGVD的基准数据集。在大规模数据集GenVidBench和HardGVD上的实验表明,该方法显著优于现有先进方法。
原文摘要 · Abstract (English)
With the rapid advancement of video generation models, distinguishing between AI-generated and authentic videos has emerged as a challenging endeavor. The majority of existing research endeavors concentrate on the development of detectors for identifying samples generated by generative adversarial networks. Nevertheless, the detection of AI-generated videos, particularly those produced by text-to-video models, still remains an uncharted territory. Although state-of-the-art text-to-video models can generate realistic visual content similar to real videos, they fall short of generating the details of the images and the changes in details within the videos. Inspired by this, we address AI-generated video detection from a novel perspective of bit-planes, which can effectively describe the details or noises in images or videos. To this end, we propose a simple yet effective approach called Noise Amplification. This approach first extracts noise signals based on bit-planes, then amplifies these noise signals, and finally feeds them into the discriminator networks for video fake classification. Noise amplification is comprehensively constructed by incorporating three aspects: pixel-level intensity enhancement, region-level spatial amplification, and frame-level temporal aggregation. To evaluate methods of AI-generated video detection in challenging scenarios, we also introduce a benchmark named HardGVD. Extensive experiments on both the large-scale dataset GenVidBench and HardGVD show that our simple approach significantly outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。