arXiv:2506.14827cs.CVcs.AI2025-06被引 8

首个可解释AI视频检测模型,能定位并说明生成缺陷。

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning

  • 用多模态标注数据训练,实现视觉推理链
  • 在多种生成器上表现良好,准确率超90%
  • 适合需要透明判断的审核与媒体平台

随着AI生成视频在媒体平台日益普及,可靠区分合成内容与真实影像变得尤为紧迫。现有方法多将其视为二分类任务,难以揭示模型判断依据。本研究提出DAVID-X,首个包含缺陷级时空标注与文字推理的AI生成视频数据集。基于此,我们构建DAVID-XR1,一种视频-语言模型,可输出包含缺陷分类、时空定位和自然语言解释的可解释推理链。该方法将黑箱检测转变为可验证的诊断流程。实验表明,仅用小规模数据微调通用骨干网络,并结合思维链蒸馏,即可在多种生成器与模式下实现强泛化能力。结果证明可解释检测对可信识别具有重要价值。

原文摘要 · Abstract (English)

As AI-generated video becomes increasingly pervasive across media platforms, the ability to reliably distinguish synthetic content from authentic footage has become both urgent and essential. Existing approaches have primarily treated this challenge as a binary classification task, offering limited insight into where or why a model identifies a video as AI-generated. However, the core challenge extends beyond simply detecting subtle artifacts; it requires providing fine-grained, persuasive evidence that can convince auditors and end-users alike. To address this critical gap, we introduce DAVID-X, the first dataset to pair AI-generated videos with detailed defect-level, temporal-spatial annotations and written rationales. Leveraging these rich annotations, we present DAVID-XR1, a video-language model designed to deliver an interpretable chain of visual reasoning-including defect categorization, temporal-spatial localization, and natural language explanations. This approach fundamentally transforms AI-generated video detection from an opaque black-box decision into a transparent and verifiable diagnostic process. We demonstrate that a general-purpose backbone, fine-tuned on our compact dataset and enhanced with chain-of-thought distillation, achieves strong generalization across a variety of generators and generation modes. Our results highlight the promise of explainable detection methods for trustworthy identification of AI-generated video content.

AI检测可解释性视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。