arXiv:2603.13739cs.CVcs.AI2026-03

统一文本与图像控制的视频生成模型,提升画面连贯性。

UniVid: Pyramid Diffusion Model for High Quality Video Generation

论文配图:UniVid: Pyramid Diffusion Model for High Quality Video Generation
图 1 · 摘自论文原文
  • 采用双流交叉注意力机制融合文本与图像信息
  • 在T2V/I2V/(T+I)2V任务中均实现更优时序连贯性
  • 支持单模态到双模态控制的自由插值,灵活可控

基于扩散模型的文本到视频(T2V)或图像到视频(I2V)生成已成为研究热点。然而,如何将两种生成范式整合到统一模型中仍面临挑战。本文提出统一视频生成模型UniVid,支持文本提示与参考图像双重控制。模型从文本中提取物体外观和运动描述,从图像中获取纹理细节与结构信息,共同引导视频生成。通过引入时序金字塔跨帧空间-时序注意力模块和卷积结构,对预训练文生图扩散模型进行扩展,以生成具有时间连贯性的视频帧。为支持双模态控制,设计双流交叉注意力机制,其注意力权重可在推理时自由重调,实现单模态与双模态控制间的平滑插值。大量实验表明,UniVid在T2V、I2V及(T+I)2V任务上均展现出优异的时序连贯性表现。

原文摘要 · Abstract (English)

Diffusion-based text-to-video generation (T2V) or image-to-video (I2V) generation have emerged as a prominent research focus. However, there exists a challenge in integrating the two generative paradigms into a unified model. In this paper, we present a unified video generation model (UniVid) with hybrid conditions of the text prompt and reference image. Given these two available controls, our model can extract objects' appearance and their motion descriptions from textual prompts, while obtaining texture details and structural information from image clues to guide the video generation process. Specifically, we scale up the pre-trained text-to-image diffusion model for generating temporally coherent frames via introducing our temporal-pyramid cross-frame spatial-temporal attention modules and convolutions. To support bimodal control, we introduce a dual-stream cross-attention mechanism, whose attention scores can be freely re-weighted for interpolation of between single and two modalities controls during inference. Extensive experiments showcase that our UniVid achieves superior temporal coherence on T2V, I2V and (T+I)2V tasks.

视频生成扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。