通过光流残差与时空一致性检测AI生成视频的细微篡改痕迹。
Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency
- 双分支结构:分别分析图像外观和光流残差,捕捉视觉与运动异常。
- 在10种生成模型上验证,对文本到视频、图像到视频任务均有效。
- 擅长发现高保真度视频中的微小时间不一致,适合真实场景检测。
扩散模型驱动的视频生成技术迅速发展,生成内容愈发逼真,给视频伪造检测带来新挑战。现有方法难以捕捉高保真度AI生成视频中精细的时序不一致。本文提出一种结合RGB外观特征与光流残差的时空一致性检测框架。模型采用双分支架构:一支分析RGB帧以识别外观层面的伪影,另一支处理光流残差以揭示由不完美时序合成引发的细微运动异常。通过融合互补特征,该方法能有效检测多种伪造视频。在涵盖十种不同生成模型的文本到视频与图像到视频任务上进行大量实验,验证了方法的鲁棒性与强泛化能力。
原文摘要 · Abstract (English)
The rapid advancement of diffusion-based video generation models has led to increasingly realistic synthetic content, presenting new challenges for video forgery detection. Existing methods often struggle to capture fine-grained temporal inconsistencies, particularly in AI-generated videos with high visual fidelity and coherent motion. In this work, we propose a detection framework that leverages spatial-temporal consistency by combining RGB appearance features with optical flow residuals. The model adopts a dual-branch architecture, where one branch analyzes RGB frames to detect appearance-level artifacts, while the other processes flow residuals to reveal subtle motion anomalies caused by imperfect temporal synthesis. By integrating these complementary features, the proposed method effectively detects a wide range of forged videos. Extensive experiments on text-to-video and image-to-video tasks across ten diverse generative models demonstrate the robustness and strong generalization ability of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。