arXiv:2507.00583cs.CVcs.AI2025-07NeurIPS被引 36

通过分析视频在神经表示中的轨迹几何特征,高效区分真实与生成视频。

AI-Generated Video Detection via Perceptual Straightening

  • 基于感知直线化假设,量化视频在视觉模型中的轨迹曲率与步长距离。
  • 在VidProM上达到97.17%准确率和98.63%AUROC,优于现有方法。
  • 轻量级设计,适合实际部署,为生成视频检测提供新视角。

生成式AI的快速发展使得合成视频日益逼真,对内容认证构成严峻挑战,并引发滥用担忧。现有检测方法常面临泛化能力差和难以捕捉细微时序不一致的问题。本文提出ReStraV(Representation Straightening Video)——一种新型自然与生成视频区分方法。受‘感知直线化’假说启发,即真实视频在神经表示空间中轨迹趋于直线,我们分析其偏离该预期几何特性的程度。利用预训练自监督视觉变压器(DINOv2),量化视频在表示域中的时序曲率与逐帧距离,并聚合每段视频的统计特征,训练分类器。分析表明,生成视频在曲率与距离模式上显著区别于真实视频。轻量级分类器在VidProM基准上实现97.17%准确率和98.63% AUROC,大幅超越现有图像与视频基方法。ReStraV计算高效,提供低成本、高效果的检测方案。本工作为利用神经表示几何特性检测生成视频提供了新思路。

原文摘要 · Abstract (English)

The rapid advancement of generative AI enables highly realistic synthetic videos, posing significant challenges for content authentication and raising urgent concerns about misuse. Existing detection methods often struggle with generalization and capturing subtle temporal inconsistencies. We propose ReStraV(Representation Straightening Video), a novel approach to distinguish natural from AI-generated videos. Inspired by the "perceptual straightening" hypothesis -- which suggests real-world video trajectories become more straight in neural representation domain -- we analyze deviations from this expected geometric property. Using a pre-trained self-supervised vision transformer (DINOv2), we quantify the temporal curvature and stepwise distance in the model's representation domain. We aggregate statistics of these measures for each video and train a classifier. Our analysis shows that AI-generated videos exhibit significantly different curvature and distance patterns compared to real videos. A lightweight classifier achieves state-of-the-art detection performance (e.g., 97.17% accuracy and 98.63% AUROC on the VidProM benchmark), substantially outperforming existing image- and video-based methods. ReStraV is computationally efficient, it is offering a low-cost and effective detection solution. This work provides new insights into using neural representation geometry for AI-generated video detection.

视频检测生成内容神经表示DINOv2

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。