用多模态大模型提升图像转视频的动态控制与连贯性。
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
- 融合多模态大模型与扩散Transformer,联合编码视觉与文本条件。
- 在DIVE评估下,动态范围、可控性和质量分别提升42.5%、7.9%、11.8%。
- 新构建的DIVE基准解决现有评测对低动态视频的偏好问题,适合复杂场景生成研究者。
近期图像到视频(I2V)生成在常规场景中表现良好,但在需深层理解细微运动和复杂物体-动作关系的复杂场景中仍面临挑战。为此,我们提出Dynamic-I2V框架,将多模态大语言模型(MLLM)融入扩散Transformer(DiT)架构,联合编码视觉与文本条件。借助MLLM的多模态理解能力,模型显著提升生成视频的运动可控性与时间连贯性。其天然多模态特性支持多样条件输入,拓展至多种下游生成任务。系统分析揭示当前I2V基准存在明显偏差:过度偏好低动态视频,源于运动复杂度与视觉质量指标的失衡。为弥合此评估缺口,我们提出DIVE——一个专为全面测量I2V动态质量而设计的新基准。大量定量与定性实验表明,Dynamic-I2V在图像到视频生成中达到领先性能,尤其在DIVE评估下,动态范围、可控性与质量分别较现有方法提升42.5%、7.9%、11.8%。
原文摘要 · Abstract (English)
Recent advancements in image-to-video (I2V) generation have shown promising performance in conventional scenarios. However, these methods still encounter significant challenges when dealing with complex scenes that require a deep understanding of nuanced motion and intricate object-action relationships. To address these challenges, we present Dynamic-I2V, an innovative framework that integrates Multimodal Large Language Models (MLLMs) to jointly encode visual and textual conditions for a diffusion transformer (DiT) architecture. By leveraging the advanced multimodal understanding capabilities of MLLMs, our model significantly improves motion controllability and temporal coherence in synthesized videos. The inherent multimodality of Dynamic-I2V further enables flexible support for diverse conditional inputs, extending its applicability to various downstream generation tasks. Through systematic analysis, we identify a critical limitation in current I2V benchmarks: a significant bias towards favoring low-dynamic videos, stemming from an inadequate balance between motion complexity and visual quality metrics. To resolve this evaluation gap, we propose DIVE - a novel assessment benchmark specifically designed for comprehensive dynamic quality measurement in I2V generation. In conclusion, extensive quantitative and qualitative experiments confirm that Dynamic-I2V attains state-of-the-art performance in image-to-video generation, particularly revealing significant improvements of 42.5%, 7.9%, and 11.8% in dynamic range, controllability, and quality, respectively, as assessed by the DIVE metric in comparison to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。