无需训练即可提升文本生成视频质量,解决动态不足与时间不一致问题。
ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
- 通过减少解码器各层的时间注意力差异,增强视频结构合理性和时间一致性。
- 利用傅里叶方法增强注意力图能量,显著提升生成视频的运动幅度和丰富度。
- 零参数新增、无额外计算开销,适合快速部署到现有T2V模型中。
文本到视频(T2V)生成模型虽具便捷视觉创作潜力,但常出现结构不合理、时间不一致及运动缺失等问题,导致近似静态视频。本文发现不同解码器块间时间注意力图的差异程度与时间不一致性相关,且注意力图能量与运动幅度直接正相关。基于此,提出训练免费的ByTheWay方法:1)时间自引导机制通过缩小各层时间注意力差异,提升视频结构合理性与时间一致性;2)基于傅里叶的运动增强通过放大注意力图能量,增强运动幅度与丰富度。大量实验表明,ByTheWay在几乎无额外开销下显著提升生成质量。
原文摘要 · Abstract (English)
The text-to-video (T2V) generation models, offering convenient visual creation, have recently garnered increasing attention. Despite their substantial potential, the generated videos may present artifacts, including structural implausibility, temporal inconsistency, and a lack of motion, often resulting in near-static video. In this work, we have identified a correlation between the disparity of temporal attention maps across different blocks and the occurrence of temporal inconsistencies. Additionally, we have observed that the energy contained within the temporal attention maps is directly related to the magnitude of motion amplitude in the generated videos. Based on these observations, we present ByTheWay, a training-free method to improve the quality of text-to-video generation without introducing additional parameters, augmenting memory or sampling time. Specifically, ByTheWay is composed of two principal components: 1) Temporal Self-Guidance improves the structural plausibility and temporal consistency of generated videos by reducing the disparity between the temporal attention maps across various decoder blocks. 2) Fourier-based Motion Enhancement enhances the magnitude and richness of motion by amplifying the energy of the map. Extensive experiments demonstrate that ByTheWay significantly improves the quality of text-to-video generation with negligible additional cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。