arXiv:2504.04051cs.CVcs.AI2025-04被引 15

测试顶级文生视频模型数数能力,发现9个以内物体几乎全错。

Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models

  • 构建人类评估基准T2VCountBench,专测文生视频计数能力。
  • 所有模型在生成9个以下物体时几乎全失败,准确率极低。
  • 适合关注视频生成可控性与指令遵循的研究者。

生成模型在各类AI任务中取得显著进展,尤其在文生视频领域,如Video LDM和Stable Video Diffusion等模型已能根据文本指令生成电影级真实视频。然而,当前文生视频模型在可靠执行人类指令方面仍存在根本性挑战,尤其是在遵循简单数量约束方面。本文提出T2VCountBench,一个面向2025年主流文生视频模型的计数能力专项评估基准。该基准采用严格的真人评估,测量生成物体数量,涵盖开源与商业模型。大量实验表明,所有现有模型在生成9个或更少物体的任务中几乎均无法达标。此外,系统性消融研究揭示视频风格、时间动态及多语言输入等因素对计数表现的影响。我们还探索提示优化技术,发现将任务拆解为子任务也无法有效缓解此缺陷。研究揭示了当前文生视频生成在数量控制方面的严重局限,并为未来提升指令遵循能力提供重要启示。

原文摘要 · Abstract (English)

Generative models have driven significant progress in a variety of AI tasks, including text-to-video generation, where models like Video LDM and Stable Video Diffusion can produce realistic, movie-level videos from textual instructions. Despite these advances, current text-to-video models still face fundamental challenges in reliably following human commands, particularly in adhering to simple numerical constraints. In this work, we present T2VCountBench, a specialized benchmark aiming at evaluating the counting capability of SOTA text-to-video models as of 2025. Our benchmark employs rigorous human evaluations to measure the number of generated objects and covers a diverse range of generators, covering both open-source and commercial models. Extensive experiments reveal that all existing models struggle with basic numerical tasks, almost always failing to generate videos with an object count of 9 or fewer. Furthermore, our comprehensive ablation studies explore how factors like video style, temporal dynamics, and multilingual inputs may influence counting performance. We also explore prompt refinement techniques and demonstrate that decomposing the task into smaller subtasks does not easily alleviate these limitations. Our findings highlight important challenges in current text-to-video generation and provide insights for future research aimed at improving adherence to basic numerical constraints.

文生视频指令遵循计数能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。