Presto模型实现15秒长视频生成,内容丰富且连贯。
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
- 采用分段交叉注意力机制,按时间分段对齐子描述文本。
- 在VBench上达78.5%语义得分,动态度100%,超越现有方法。
- 适用于需要长时序一致性和高内容密度的视频生成任务。
我们提出Presto,一种新型视频扩散模型,可生成15秒长、具有长期连贯性与丰富内容的视频。将视频生成扩展至长时间跨度以维持场景多样性面临重大挑战。为此,我们提出分段交叉注意力(SCA)策略:沿时间维度将隐藏状态分段,使每段分别对对应子描述进行交叉注意力。SCA无需额外参数,可无缝集成到当前基于DiT的架构中。为支持高质量长视频生成,我们构建了LongTake-HD数据集,包含26.1万条内容丰富的视频,具备场景一致性,并配有整体视频描述及五个渐进式子描述。实验表明,Presto在VBench语义得分上达到78.5%,动态度达100%,显著优于现有先进方法。结果证明Presto有效提升了内容丰富性,保持了长期连贯性,并精准捕捉复杂文本细节。更多细节见项目页:https://presto-video.github.io/。
原文摘要 · Abstract (English)
We introduce Presto, a novel video diffusion model designed to generate 15-second videos with long-range coherence and rich content. Extending video generation methods to maintain scenario diversity over long durations presents significant challenges. To address this, we propose a Segmented Cross-Attention (SCA) strategy, which splits hidden states into segments along the temporal dimension, allowing each segment to cross-attend to a corresponding sub-caption. SCA requires no additional parameters, enabling seamless incorporation into current DiT-based architectures. To facilitate high-quality long video generation, we build the LongTake-HD dataset, consisting of 261k content-rich videos with scenario coherence, annotated with an overall video caption and five progressive sub-captions. Experiments show that our Presto achieves 78.5% on the VBench Semantic Score and 100% on the Dynamic Degree, outperforming existing state-of-the-art video generation methods. This demonstrates that our proposed Presto significantly enhances content richness, maintains long-range coherence, and captures intricate textual details. More details are displayed on our project page: https://presto-video.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。