arXiv:2510.26412cs.CVcs.AI2025-10中稿 · ICML被引 5

构建首个长视频文本生成评测基准,揭示模型在角色一致性上的短板。

LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation

  • 基于真实视频构建多场景复杂文本提示的长视频评测集
  • 17个模型测试显示感知质量高但角色和细节对齐弱
  • 适合关注长视频生成真实性和连贯性的研究者使用

近期文本到视频生成在短片段上取得显著进展,但在复杂文本输入下的长视频生成评估仍具挑战。为此,我们提出LoCoT2V-Bench,一个面向长视频生成(LVG)的基准,包含基于真实视频构建的多场景提示及层级化元数据(如角色设定、摄像机行为)。同时提出多维评估框架LoCoT2V-Eval,涵盖感知质量、文本-视频对齐、时间质量、动态质量及人类期望实现度(HERD),特别关注细粒度对齐与角色时序一致性。在17个代表性LVG模型上的实验表明,各维度能力差异明显:感知质量与背景一致性较强,但细粒度对齐与角色一致性显著不足。结果表明,提升提示忠实度与身份保持仍是长视频生成的核心挑战。代码与数据已开源。

原文摘要 · Abstract (English)

Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present LoCoT2V-Bench, a benchmark for long video generation (LVG) featuring multi-scene prompts with hierarchical metadata (e.g., character settings and camera behaviors), constructed from collected real-world videos. We further propose LoCoT2V-Eval, a multi-dimensional framework covering perceptual quality, text-video alignment, temporal quality, dynamic quality, and Human Expectation Realization Degree (HERD), with an emphasis on aspects such as fine-grained text-video alignment and temporal character consistency. Experiments on 17 representative LVG models reveal pronounced capability disparities across evaluation dimensions, with strong perceptual quality and background consistency but markedly weaker fine-grained text-video alignment and character consistency. These findings suggest that improving prompt faithfulness and identity preservation remains a key challenge for long-form video generation. Our code and data are released at https://github.com/XqZeppelinhead0702/LoCoT2V-Bench

文本生成视频长视频生成评测基准角色一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。