arXiv:2511.05573cs.CVcs.AI2025-11被引 1

用合成视频提升文本生成清晰度,轻量高效。

Video Text Preservation with Synthetic Text-Rich Videos

  • 用文生图+无文本图生视频生成合成数据,用于微调。
  • 微调后短文本可读性提升,长文本结构更连贯。
  • 无需改模型架构,适合快速部署到实际应用。

尽管文生视频(T2V)模型发展迅速,但在生成视频中可读且连贯的文本方面仍存在困难,尤其对短语或单词的渲染效果不佳。现有解决方法计算成本高,不适用于视频生成。本文提出一种轻量级方法:先用文生图扩散模型生成富文本图像,再通过无文本的图生视频模型将其动画化为短视频,形成合成视频-提示对,用于微调预训练的T2V模型Wan2.1,不改变模型结构。实验表明,该方法提升了短文本的可读性与长文本的时间一致性,表明经过筛选的合成数据与弱监督是实现高质量文本保真生成的可行路径。

原文摘要 · Abstract (English)

While Text-To-Video (T2V) models have advanced rapidly, they continue to struggle with generating legible and coherent text within videos. In particular, existing models often fail to render correctly even short phrases or words and previous attempts to address this problem are computationally expensive and not suitable for video generation. In this work, we investigate a lightweight approach to improve T2V diffusion models using synthetic supervision. We first generate text-rich images using a text-to-image (T2I) diffusion model, then animate them into short videos using a text-agnostic image-to-video (I2v) model. These synthetic video-prompt pairs are used to fine-tune Wan2.1, a pre-trained T2V model, without any architectural changes. Our results show improvement in short-text legibility and temporal consistency with emerging structural priors for longer text. These findings suggest that curated synthetic data and weak supervision offer a practical path toward improving textual fidelity in T2V generation.

文生视频文本生成扩散模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。