arXiv:2506.20103cs.CVcs.AI2025-06被引 3

构建首个像素级视频伪影定位数据集,助力AI生成视频质量评估

BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos

  • 构建3254个生成视频的像素级伪影标注数据集
  • 训练模型在该数据集上定位精度显著提升
  • 适合研究生成视频质量控制与模型优化的学者

深度生成模型在视频生成方面取得进展,但合成内容仍存在时序运动不一致、物理轨迹不合理、物体变形不自然和局部模糊等视觉伪影,影响真实感与用户信任。准确检测并定位这些伪影对自动化质量控制和改进生成模型至关重要。然而,当前缺乏专门用于生成视频伪影定位的全面基准数据集。现有数据集或仅支持视频/帧级别检测,或缺少精细的空间标注。为此,我们提出BrokenVideos,包含3,254个AI生成视频,配有经人工仔细验证的像素级掩码,精准标注视觉缺陷区域。实验表明,在BrokenVideos上训练的先进伪影检测模型及多模态大语言模型(MLLMs)在定位能力上显著增强。通过广泛评估,证明BrokenVideos为生成视频伪影定位研究提供了关键基准。数据集地址:https://broken-video-detection-datetsets.github.io/Broken-Video-Detection-Datasets.github.io/

原文摘要 · Abstract (English)

Recent advances in deep generative models have led to significant progress in video generation, yet the fidelity of AI-generated videos remains limited. Synthesized content often exhibits visual artifacts such as temporally inconsistent motion, physically implausible trajectories, unnatural object deformations, and local blurring that undermine realism and user trust. Accurate detection and spatial localization of these artifacts are crucial for both automated quality control and for guiding the development of improved generative models. However, the research community currently lacks a comprehensive benchmark specifically designed for artifact localization in AI generated videos. Existing datasets either restrict themselves to video or frame level detection or lack the fine-grained spatial annotations necessary for evaluating localization methods. To address this gap, we introduce BrokenVideos, a benchmark dataset of 3,254 AI-generated videos with meticulously annotated, pixel-level masks highlighting regions of visual corruption. Each annotation is validated through detailed human inspection to ensure high quality ground truth. Our experiments show that training state of the art artifact detection models and multi modal large language models (MLLMs) on BrokenVideos significantly improves their ability to localize corrupted regions. Through extensive evaluation, we demonstrate that BrokenVideos establishes a critical foundation for benchmarking and advancing research on artifact localization in generative video models. The dataset is available at: https://broken-video-detection-datetsets.github.io/Broken-Video-Detection-Datasets.github.io/.

视频生成伪影检测数据集质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。