从检测视频瑕疵转向验证内容是否符合现实事实
Detecting AI-Generated Video: A Vision-Language Dual-View Survey

- 提出视觉语言双视角分类体系,分四层组织检测方法
- 基于221篇论文系统梳理生成范式与评估标准
- 适合关注可信视频检测、多模态推理的研究者
AI生成视频(AIGC-V)的逼真度不断提升,传统依赖低级伪影的检测手段已显不足,亟需从低层次检查转向高层次语义验证。本文将AIGC-V检测重新定义为事实一致性验证,即判断视频中事件、实体和物理过程是否符合现实。为系统化这一快速发展的领域,提出视觉-语言双视角分类框架,构建涵盖内在线索分析、时空一致性建模、跨模态一致性推理与语言引导的世界级推理四个层级的层次化体系。该框架凸显了从传统深度伪造检测中的伪影匹配,向依托视觉语言模型与智能体推理流程的证据驱动语义验证的根本转变。基于对221项工作的系统综述,本文总结了AIGC-V生成范式,调研了检测方法全景,并回顾了与所提视角一致的评估指标与基准数据集。最后,探讨当前挑战并指明鲁棒性、可解释性与可信检测的未来方向。
原文摘要 · Abstract (English)
The evolving realism of AI-generated Videos (AIGC-V) is rapidly rendering traditional artifact-centric detection insufficient, necessitating a paradigm shift from low-level inspection to high-level semantic verification. This paper presents a comprehensive survey of AIGC-V detection, reframing the task as Factual Fidelity Verification, which asks whether the events, entities, and physical processes depicted in a video are consistent with real-world facts. To systematize this rapidly evolving field, we propose a Vision-Language Dual-View taxonomy that organizes existing methods into a hierarchical, four-layer landscape, spanning intrinsic cue analysis, spatiotemporal consistency modeling, cross-modal consistency reasoning, and language-guided world-level reasoning. This dual-view framing highlights a fundamental transition from artifact matching in traditional deepfake detection to evidence-based semantic verification enabled by vision-language models and agentic reasoning pipelines. Based on a systematic review of 221 works, we synthesize AIGC-V generation paradigms, survey the landscape of detection methods, and review evaluation metrics and benchmarks in line with proposed views. Finally, we discuss current challenges and identify promising directions toward robust, explainable, and trustworthy detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。