发现运动类视频伪造检测模型依赖数据偏见,真实场景下效果大幅下降。
Dataset Biases and Shortcut Learning in Motion-Based AI-Generated Video Detection

- 检测模型依赖训练数据中的运动模式偏差
- 换数据集后准确率暴跌至随机水平
- 频域方法更稳定,适合通用检测
近年来,AI生成视频的视觉质量显著提升,使人难以分辨真伪。本文评估了四种先进运动特征检测器的鲁棒性与适用性,发现其性能严重依赖于训练数据中的预处理和采样偏差。这些检测器对特定数据集中的运动模式高度敏感——通常生成视频的帧间运动少于真实视频。当在不含此类运动偏见的数据集上测试时,所有检测器性能均降至接近随机水平。通过数据集重平衡和简单空间增强后,所有模型性能均出现严重退化。相比之下,现有频域检测器在所有数据集上保持强性能,表明频域方法可能更具泛化能力。本文呼吁关注此类漏洞,推动更代表性和无偏的评估体系发展。
原文摘要 · Abstract (English)
The visual quality of AI-generated videos has improved drastically in recent years, making it increasingly difficult for humans to distinguish between real and synthetic media. In this work, we evaluate the robustness and applicability of four state-of-the-art motion-based AI-generated video detectors. We identify significant preprocessing and sampling biases in these methods and demonstrate that they account for a substantial portion of their reported performance. Furthermore, we find that these detectors are highly sensitive to motion patterns specific to their evaluation datasets, where AI-generated videos generally exhibit less inter-frame movement than real videos. We show that for all detectors, performance collapses to near-random levels when evaluated on a dataset that does not contain this motion bias. Additionally, through dataset rebalancing and the application of simple spatial augmentations, we observe severe performance degradation across all evaluated models. In contrast, we find that an existing frequency-based detector maintains strong performance across all evaluated datasets, suggesting that frequency-based approaches may offer a more generalizable path forward for AI-generated video detection. We hope that our work raises awareness towards these vulnerabilities and encourages the development of more representative, unbiased datasets and more robust evaluation protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。