arXiv:2510.07441cs.CV2025-10被引 1

针对动态镜头视频生成评估缺失,提出新基准DynamicEval

DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis

  • 构建聚焦动态镜头的文本到视频评估基准,含45k人工标注
  • 提出背景与前景一致性新度量,比现有方法提升2%以上相关性
  • 适合关注视频生成质量评估、尤其动态镜头场景的研究者

现有文本到视频(T2V)评估基准如VBench和EvalCrafter存在两大局限:(i) 侧重主体为中心的提示或静态摄像机场景,对产生电影级镜头至关重要的动态镜头缺乏评估,现有指标在动态运动下研究不足;(ii) 通常将视频级评分聚合为单一模型级分数用于排名,忽略了对具体视频的评估,而这对从同一提示生成的候选视频中选择更优结果至关重要。为弥补这些空白,我们提出DynamicEval,一个系统化构建的基准,包含强调动态摄像机运动的提示,并配有来自10个T2V模型生成的3000个视频中45,000条视频对的人工标注。DynamicEval评估两个关键维度的视频质量:背景场景一致性和前景物体一致性。对于背景一致性,基于Vbench运动平滑度指标生成可解释的误差图。我们发现,尽管该指标与人类判断具有较好对齐,但在摄像机与前景物体运动导致遮挡/非遮挡时失效。据此,我们提出一种新背景一致性度量,利用物体误差图以合理方式修正这两类失败情况。第二个创新是引入前景一致性度量,通过跟踪每个物体实例内点及其邻域来评估物体保真度。大量实验表明,所提度量在视频级和模型级均与人类偏好有更强相关性(提升超过2个百分点),确立DynamicEval作为动态镜头下T2V模型评估的更全面基准。

原文摘要 · Abstract (English)

Existing text-to-video (T2V) evaluation benchmarks, such as VBench and EvalCrafter, suffer from two limitations. (i) While the emphasis is on subject-centric prompts or static camera scenes, camera motion essential for producing cinematic shots and existing metrics under dynamic motion are largely unexplored. (ii) These benchmarks typically aggregate video-level scores into a single model-level score for ranking generative models. Such aggregation, however, overlook video-level evaluation, which is vital to selecting the better video among the candidate videos generated for a given prompt. To address these gaps, we introduce DynamicEval, a benchmark consisting of systematically curated prompts emphasizing dynamic camera motion, paired with 45k human annotations on video pairs from 3k videos generated by ten T2V models. DynamicEval evaluates two key dimensions of video quality: background scene consistency and foreground object consistency. For background scene consistency, we obtain the interpretable error maps based on the Vbench motion smoothness metric. We observe that while the Vbench motion smoothness metric shows promising alignment with human judgments, it fails in two cases: occlusions/disocclusions arising from camera and foreground object movements. Building on this, we propose a new background consistency metric that leverages object error maps to correct two failure cases in a principled manner. Our second innovation is the introduction of a foreground consistency metric that tracks points and their neighbors within each object instance to assess object fidelity. Extensive experiments demonstrate that our proposed metrics achieve stronger correlations with human preferences at both the video level and the model level (an improvement of more than 2% points), establishing DynamicEval as a more comprehensive benchmark for evaluating T2V models under dynamic camera motion.

视频生成评估基准动态镜头一致性度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。