arXiv:2605.21977cs.CVcs.AI2026-05被引 1

用视频帧提升图像与视频生成内容检测的统一性

Video as Natural Augmentation: Towards Unified AI-Generated Image and Video Detection

论文配图:Video as Natural Augmentation: Towards Unified AI-Generated Image and Video Detection
图 1 · 摘自论文原文
  • 将视频帧作为自然增强数据,联合训练图像与视频检测
  • 在14个基准上实现跨模态性能最优,且无需额外调参
  • 适合需要统一检测生成内容的研究者与安全团队

AI生成内容(AIGC)快速演进,亟需能跨数据源、部署管道和视觉模态泛化的检测器。当前最先进的图像生成内容检测器在处理视频帧时往往失效。我们系统分析发现,这种跨模态差距源于视频处理中的合成无关偏移(如色彩转换、编码压缩、缩放、模糊)以及现代视频生成器引入的模型特异性指纹。为此,我们提出VINA(Video as Natural Augmentation)框架,联合训练图像与视频数据。VINA将视频帧视为物理真实的自然增强,并引入跨模态监督对比学习,对齐图像与视频表征,共享真实/虚假决策边界。在14个图像、视频及真实世界基准上的实验表明,VINA实现双向增益,显著提升鲁棒性与可迁移性,在几乎所有测试场景中达到最佳性能,且无需复杂增强或数据集特定调优。

原文摘要 · Abstract (English)

AI-generated content (AIGC) is rapidly improving, creating an urgent need for detectors that generalize across data sources, deployment pipelines, and visual modalities. A strongly generalizable detector should remain robust under distributional variations. However, we identify a consistent failure mode: SOTA AI-generated image detectors often collapse when applied to frames extracted from videos. Through systematic analysis, we show that this cross-modal gap arises from both entangled synthesis-agnostic video processing shifts, including color conversion, codec compression, resizing, and blur, and model-specific fingerprints introduced by modern video generators. Motivated by these findings, we propose VINA (Video as Natural Augmentation), a unified AIGC detection framework that jointly trains on image and video data. VINA uses video frames as physically grounded natural augmentations and further introduces a cross-modal supervised contrastive objective to align image and video representations under a shared real/fake decision boundary. Extensive experiments on 14 image, video, and in-the-wild benchmarks show that VINA delivers bidirectional gains, improves robustness and transferability, and achieves state-of-the-art performance across nearly all evaluated settings without complex augmentation or dataset-specific tuning.

生成内容检测视频生成跨模态统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。