新基准评估图像生成视频的语义理解与推理能力
UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark
- 用多模态大模型构建细粒度语义评估流程
- 涵盖空间、属性、类别和因果推理四维度,共500组图文对
- 兼顾机器与人工评价,适合研究视频生成理解力的学者
生成式扩散模型发展迅速,图像到视频(I2V)生成成为视频合成领域的重点。然而,现有评估基准主要关注视频质量与时间一致性,忽视了模型对输入图像中特定主体语义的理解能力,以及生成视频是否符合物理规律和人类常识。为此,我们提出UI2V-Bench,一个聚焦语义理解与推理的I2V生成评估基准。该基准引入四个核心评价维度:空间理解、属性绑定、类别理解与推理。为评估这些维度,我们设计基于多模态大语言模型(MLLMs)的两种方法:实例级语义理解流水线,以及支持逐步因果分析的反馈式推理流水线。UI2V-Bench包含约500组精心构造的图文对,覆盖多种开源与闭源I2V模型在所有定义维度上的评估。进一步的人工评估显示,与所提出的MLLM指标高度一致。总体而言,UI2V-Bench填补了I2V评估中对语义理解与推理能力关注的空白,提供了一个稳健的框架与数据集,助力未来相关研究与模型发展。
原文摘要 · Abstract (English)
Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing evaluation benchmarks primarily focus on aspects such as video quality and temporal consistency, while largely overlooking the model's ability to understand the semantics of specific subjects in the input image or to ensure that the generated video aligns with physical laws and human commonsense. To address this gap, we propose UI2V-Bench, a novel benchmark for evaluating I2V models with a focus on semantic understanding and reasoning. It introduces four primary evaluation dimensions: spatial understanding, attribute binding, category understanding, and reasoning. To assess these dimensions, we design two evaluation methods based on Multimodal Large Language Models (MLLMs): an instance-level pipeline for fine-grained semantic understanding, and a feedback-based reasoning pipeline that enables step-by-step causal assessment for more accurate evaluation. UI2V-Bench includes approximately 500 carefully constructed text-image pairs and evaluates a range of both open source and closed-source I2V models across all defined dimensions. We further incorporate human evaluations, which show strong alignment with the proposed MLLM-based metrics. Overall, UI2V-Bench fills a critical gap in I2V evaluation by emphasizing semantic comprehension and reasoning ability, offering a robust framework and dataset to support future research and model development in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。