用细粒度问答评估文本与视频语义对齐,更贴近人类判断。
ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering
- 通过多智能体生成原子级问题,精细分解文本语义
- 结合常识知识与多阶段推理,回答准确率达58.47相关性
- 适合评估文本生成视频模型的细节对齐能力
文本到视频生成中的语义对齐精准评估仍是挑战。现有指标如CLIPScore仅提供粗粒度分数,无法反映人类偏好。为此,我们提出ETVA:基于细粒度问题生成与回答的文本-视频对齐评估方法。首先,多智能体系统将提示解析为语义场景图并生成原子问题;随后设计知识增强的多阶段推理框架,先由辅助大模型检索常识知识(如物理规律),再由视频大模型通过多阶段推理作答。大量实验表明,ETVA在人类判断上的斯皮尔曼相关系数达58.47,显著高于现有指标的31.0。我们还构建了涵盖2000个多样化提示和12000个原子问题的基准数据集,覆盖10个类别。通过对15个现有T2V模型的系统评估,揭示其核心能力与局限,为下一代T2V生成提供方向。
原文摘要 · Abstract (English)
Precisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without fine-grained alignment details, failing to align with human preference. To address this limitation, we propose ETVA, a novel Evaluation method of Text-to-Video Alignment via fine-grained question generation and answering. First, a multi-agent system parses prompts into semantic scene graphs to generate atomic questions. Then we design a knowledge-augmented multi-stage reasoning framework for question answering, where an auxiliary LLM first retrieves relevant common-sense knowledge (e.g., physical laws), and then video LLM answers the generated questions through a multi-stage reasoning mechanism. Extensive experiments demonstrate that ETVA achieves a Spearman's correlation coefficient of 58.47, showing a much higher correlation with human judgment than existing metrics which attain only 31.0. We also construct a comprehensive benchmark specifically designed for text-to-video alignment evaluation, featuring 2k diverse prompts and 12k atomic questions spanning 10 categories. Through a systematic evaluation of 15 existing text-to-video models, we identify their key capabilities and limitations, paving the way for next-generation T2V generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。