arXiv:2505.04946cs.CVcs.AI2025-05被引 11

首个评估视频生成中文字准确性的真人评测基准

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models

  • 设计包含动态场景的复杂文本提示,测试模型跨帧保持文本一致的能力
  • 对10个主流模型测评发现,多数无法生成清晰连贯的文字
  • 适合关注视频生成精准性与真实应用落地的研究者

得益于可扩展深度架构和大规模预训练的进展,文本到视频生成在高保真度、指令遵循内容生成方面取得了前所未有的能力,广泛应用于广告、娱乐和教育。然而,这些模型在呈现精确屏幕文字(如字幕或数学公式)方面的表现尚未充分验证,这给需要文字准确性的应用带来了挑战。本文提出T2VTextBench,首个专注于评估文本到视频模型中屏幕文字保真度与时间一致性的真人评测基准。该基准通过整合复杂文本字符串与动态场景变化的提示,测试模型在多帧间维持详细指令的能力。我们评估了10个最先进的系统,涵盖开源与商业方案,结果表明大多数模型在生成可读且一致的文字方面存在明显困难。这一发现揭示了当前视频生成模型在文本操控上的关键短板,并为未来研究指明了方向。

原文摘要 · Abstract (English)

Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelity, instruction-following content across a wide range of styles, enabling applications in advertising, entertainment, and education. However, these models' ability to render precise on-screen text, such as captions or mathematical formulas, remains largely untested, posing significant challenges for applications requiring exact textual accuracy. In this work, we introduce T2VTextBench, the first human-evaluation benchmark dedicated to evaluating on-screen text fidelity and temporal consistency in text-to-video models. Our suite of prompts integrates complex text strings with dynamic scene changes, testing each model's ability to maintain detailed instructions across frames. We evaluate ten state-of-the-art systems, ranging from open-source solutions to commercial offerings, and find that most struggle to generate legible, consistent text. These results highlight a critical gap in current video generators and provide a clear direction for future research aimed at enhancing textual manipulation in video synthesis.

视频生成文本控制人类评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。