arXiv:2601.20297cs.CV2026-01被引 4

针对视频生成中的瑕疵问题,提出细粒度检测与分类方法

Artifact-Aware Evaluation for High-Quality Video Generation

  • 构建10类常见生成瑕疵的分类体系,覆盖外观、运动、镜头三方面
  • 创建8万条标注视频数据集GenVID,支持瑕疵精准定位与识别
  • 开发DVAR框架,实现对低质内容的有效过滤,适合质量评估场景

随着视频生成技术的快速发展,生成视频的评估与审计日益重要。现有方法多提供粗粒度质量评分,缺乏对具体瑕疵的定位与分类。本文提出一个聚焦外观、运动和镜头三方面影响人类感知的综合评估协议,定义了10类反映常见生成失败的瑕疵类别。为实现鲁棒的瑕疵检测与分类,构建了大规模数据集GenVID,包含由多种顶尖视频生成模型生成的8万条视频,每条均针对上述瑕疵类别进行细致标注。基于GenVID,开发DVAR框架,实现生成瑕疵的细粒度识别与分类。大量实验表明,该方法显著提升瑕疵检测准确率,并有效支持低质量内容过滤。

原文摘要 · Abstract (English)

With the rapid advancement of video generation techniques, evaluating and auditing generated videos has become increasingly crucial. Existing approaches typically offer coarse video quality scores, lacking detailed localization and categorization of specific artifacts. In this work, we introduce a comprehensive evaluation protocol focusing on three key aspects affecting human perception: Appearance, Motion, and Camera. We define these axes through a taxonomy of 10 prevalent artifact categories reflecting common generative failures observed in video generation. To enable robust artifact detection and categorization, we introduce GenVID, a large-scale dataset of 80k videos generated by various state-of-the-art video generation models, each carefully annotated for the defined artifact categories. Leveraging GenVID, we develop DVAR, a Dense Video Artifact Recognition framework for fine-grained identification and classification of generative artifacts. Extensive experiments show that our approach significantly improves artifact detection accuracy and enables effective filtering of low-quality content.

视频生成瑕疵检测质量评估数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。