arXiv:2501.09982cs.CVcs.AI2025-01被引 8

通过嵌入空间插值增强文本到视频的提示能力

RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation

  • 在文本嵌入空间中插值寻找最优文本表示
  • 可生成具有复杂特征的视频,提升生成质量
  • 适合需要精细控制视频内容的研究者

文本到视频生成模型虽取得显著进展,但仍难以生成具有复杂特征的视频。这一局限常源于文本编码器无法生成准确的嵌入,从而阻碍视频生成模型。本文提出一种新方法:通过在嵌入空间中插值选择最优文本嵌入。实验表明,该方法能使视频生成模型产出期望视频。此外,我们设计了一种基于垂足嵌入与余弦相似度的简单算法,用于识别最优插值嵌入。研究强调了准确文本嵌入的重要性,并为提升文本到视频生成性能提供了新路径。

原文摘要 · Abstract (English)

Text-to-video generation models have made impressive progress, but they still struggle with generating videos with complex features. This limitation often arises from the inability of the text encoder to produce accurate embeddings, which hinders the video generation model. In this work, we propose a novel approach to overcome this challenge by selecting the optimal text embedding through interpolation in the embedding space. We demonstrate that this method enables the video generation model to produce the desired videos. Additionally, we introduce a simple algorithm using perpendicular foot embeddings and cosine similarity to identify the optimal interpolation embedding. Our findings highlight the importance of accurate text embeddings and offer a pathway for improving text-to-video generation performance.

文本生成视频生成嵌入空间插值优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。