通过3D几何一致性检测AI生成视频,准确率显著优于现有方法。
Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency

- 利用消失点捕捉视频3D几何特征,揭示真实与生成视频差异。
- 在静态场景数据集上实现98.7%准确率,跨域泛化能力强。
- 提出几何感知注意力机制,适合视频真伪检测研究者使用。
基于扩散模型的生成技术已能产出高度逼真的视频,亟需可靠的检测手段。然而,现有方法对生成视频中蕴含的3D几何模式探索有限。本文以消失点作为3D几何模式的显式表征,揭示了真实视频与AI生成视频在几何一致性上的根本差异。我们提出Grab-3D,一种基于3D几何时序一致性的生成视频检测框架。为实现可靠评估,构建了一个静态场景的AI生成视频数据集,支持稳定的3D几何特征提取。该框架采用几何感知变换器,包含几何位置编码、时序-几何注意力机制及基于EMA的几何分类头,显式注入3D几何先验知识。实验表明,Grab-3D显著优于现有最先进检测器,在未见生成器上仍保持鲁棒的跨域泛化能力。
原文摘要 · Abstract (English)
Recent advances in diffusion-based generation techniques enable AI models to produce highly realistic videos, heightening the need for reliable detection mechanisms. However, existing detection methods provide only limited exploration of the 3D geometric patterns present in generated videos. In this paper, we use vanishing points as an explicit representation of 3D geometry patterns, revealing fundamental discrepancies in geometric consistency between real and AI-generated videos. We introduce Grab-3D, a geometry-aware transformer framework for detecting AI-generated videos based on 3D geometric temporal consistency. To enable reliable evaluation, we construct an AI-generated video dataset of static scenes, allowing stable 3D geometric feature extraction. We propose a geometry-aware transformer equipped with geometric positional encoding, temporal-geometric attention, and an EMA-based geometric classifier head to explicitly inject 3D geometric awareness into temporal modeling. Experiments demonstrate that Grab-3D significantly outperforms state-of-the-art detectors, achieving robust cross-domain generalization to unseen generators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。