Skyra通过分析视觉伪影实现可解释的AI生成视频检测
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- 基于细粒度伪影识别构建多模态模型
- 在3000个样本上准确率超越现有方法
- 适合需要可信检测结果的审核与安全场景
AI视频生成技术的滥用引发严重社会担忧,亟需可靠的检测手段。现有方法多为二分类且缺乏可解释性。本文提出Skyra,一种专用于识别人类可感知视觉伪影的多模态大模型,利用这些伪影作为可解释证据进行检测。我们构建了首个大规模带细粒度人工标注的AI生成视频伪影数据集ViF-CoT-4K,用于监督微调。采用两阶段训练策略,系统提升模型在时空伪影感知、解释能力与检测精度方面表现。为全面评估,我们引入包含3000个高质量样本的ViF-Bench基准,覆盖十余种顶尖视频生成器。大量实验表明,Skyra在多个基准上均优于现有方法,评估结果为可解释AI生成视频检测的发展提供了重要洞见。
原文摘要 · Abstract (English)
The misuse of AI-driven video generation technologies has raised serious social concerns, highlighting the urgent need for reliable AI-generated video detectors. However, most existing methods are limited to binary classification and lack the necessary explanations for human interpretation. In this paper, we present Skyra, a specialized multimodal large language model (MLLM) that identifies human-perceivable visual artifacts in AI-generated videos and leverages them as grounded evidence for both detection and explanation. To support this objective, we construct ViF-CoT-4K for Supervised Fine-Tuning (SFT), which represents the first large-scale AI-generated video artifact dataset with fine-grained human annotations. We then develop a two-stage training strategy that systematically enhances our model's spatio-temporal artifact perception, explanation capability, and detection accuracy. To comprehensively evaluate Skyra, we introduce ViF-Bench, a benchmark comprising 3K high-quality samples generated by over ten state-of-the-art video generators. Extensive experiments demonstrate that Skyra surpasses existing methods across multiple benchmarks, while our evaluation yields valuable insights for advancing explainable AI-generated video detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。