评测大模型识别AI视频伪影的能力,发现多数表现接近随机
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

- 构建三层伪影分类体系,覆盖写实、动画、CG三类视频
- 19个大模型在复杂场景下识别伪影准确率接近随机
- 大模型判断与人类偏好严重不符,可靠性存疑
近期视频生成模型虽显著提升真实感,但仍存在时间不一致、结构扭曲和语义断裂等伪影。尽管多模态大模型具备强视觉理解能力,其对这类伪影的感知与推理能力尚不明确。现有基准普遍缺乏对伪影感知的系统评估及细粒度诊断推理能力的检验,尤其在超越写实内容的多样化生成视频领域。为此,我们提出Artifact-Bench,一个全面评估多模态大模型在AI生成视频伪影检测与分析方面的基准。首先建立包含写实、动画、CG风格的三层伪影层级分类体系,据此定义三项互补任务:真实与AI生成视频分类、成对真实感对比、细粒度伪影识别。在19个领先多模态大模型上的实验表明,其在伪影感知与推理方面存在显著局限,许多模型在挑战性场景中表现接近甚至低于随机水平。进一步观察到大模型判断与人类感知偏好存在显著偏差,揭示其作为通用真实性评估工具的可靠性有限。
原文摘要 · Abstract (English)
Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While Multimodal Large Language Models (MLLMs) show strong visual understanding capabilities, their ability to perceive and reason about such artifacts remains unclear. Existing benchmarks often lack systematic evaluation of artifact-aware perception and fine-grained diagnostic reasoning, especially across diverse AI-generated video domains beyond photorealistic content. To address this gap, we introduce Artifact-Bench, a comprehensive benchmark for evaluating MLLMs on AI-generated video artifact detection and analysis. We first establish a three-level hierarchical taxonomy of realism artifacts, covering photorealistic, animated, and CG-style videos. Based on this taxonomy, Artifact-Bench defines three complementary tasks: real vs. AI-generated video classification, pairwise realism comparison, and fine-grained artifact identification. Experiments on 19 leading MLLMs reveal substantial limitations in artifact perception and reasoning, with many models approaching random or even below-random performance in challenging settings. We further observe significant misalignment between MLLM judgments and human perceptual preferences, highlighting their limited reliability as general evaluators for AI-generated video realism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。