用视觉Transformer检测假视频,效果好且适应性强。
Advance Fake Video Detection via Vision Transformers
- 将视觉Transformer的时序嵌入融合用于假视频检测
- 在多个生成方法数据集上准确率高,支持少样本学习
- 适合需要通用检测能力的AI安全研究者
近年来,基于AI的多媒体生成技术已能制作高度逼真的图像和视频,引发虚假信息传播担忧。生成技术的广泛可用性和持续改进,凸显了对高精度、强泛化能力的AI生成媒体检测方法的迫切需求,这也受到欧盟《数字人工智能法案》等新规的推动。本文受基于视觉Transformer的假图像检测启发,将其拓展至视频领域。提出一种原创性框架,通过有效整合视觉Transformer的时序嵌入,提升检测性能。该方法在包含五种主流开源生成技术及专有生成方法的大型多样化视频数据集上,展现出优异的准确性、泛化能力与少样本学习潜力。
原文摘要 · Abstract (English)
Recent advancements in AI-based multimedia generation have enabled the creation of hyper-realistic images and videos, raising concerns about their potential use in spreading misinformation. The widespread accessibility of generative techniques, which allow for the production of fake multimedia from prompts or existing media, along with their continuous refinement, underscores the urgent need for highly accurate and generalizable AI-generated media detection methods, underlined also by new regulations like the European Digital AI Act. In this paper, we draw inspiration from Vision Transformer (ViT)-based fake image detection and extend this idea to video. We propose an {original} %innovative framework that effectively integrates ViT embeddings over time to enhance detection performance. Our method shows promising accuracy, generalization, and few-shot learning capabilities across a new, large and diverse dataset of videos generated using five open source generative techniques from the state-of-the-art, as well as a separate dataset containing videos produced by proprietary generative methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。