arXiv:2511.00107cs.CVcs.AI2025-11

MOVAI提升文本生成视频的连贯性与细节质量,支持复杂场景精准控制。

AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency

  • 分层解析文本为带时间标注的场景图,实现语义与时序理解
  • 通过时空注意力机制保持动作连贯,多尺度细化提升画质
  • 适合需要精细叙事控制和动态一致性的视频生成任务

文本到视频生成是生成式人工智能的关键前沿,但现有方法在时间一致性、场景理解与视觉叙事控制方面仍存在不足。本文提出MOVAI(Multimodal Original Video AI),一种融合组合场景理解与时序感知扩散模型的分层框架,用于高质量文本到视频合成。核心创新包括:(1) 组合场景解析器(CSP),将文本描述分解为带时间标注的层次化场景图;(2) 时空注意力机制(TSAM),在保持空间细节的同时确保帧间运动连贯性;(3) 渐进式视频精炼模块(PVR),通过多尺度时序推理迭代优化视频质量。在标准基准上的大量实验表明,MOVAI在各项指标上达到领先水平:相比现有方法,LPIPS提升15.3%,FVD提升12.7%,用户偏好研究中提升18.9%。该框架在生成多物体复杂场景、具备真实时间动态与细粒度语义控制方面表现突出。

原文摘要 · Abstract (English)

Text to video generation has emerged as a critical frontier in generative artificial intelligence, yet existing approaches struggle with maintaining temporal consistency, compositional understanding, and fine grained control over visual narratives. We present MOVAI (Multimodal Original Video AI), a novel hierarchical framework that integrates compositional scene understanding with temporal aware diffusion models for high fidelity text to video synthesis. Our approach introduces three key innovations: (1) a Compositional Scene Parser (CSP) that decomposes textual descriptions into hierarchical scene graphs with temporal annotations, (2) a Temporal-Spatial Attention Mechanism (TSAM) that ensures coherent motion dynamics across frames while preserving spatial details, and (3) a Progressive Video Refinement (PVR) module that iteratively enhances video quality through multi-scale temporal reasoning. Extensive experiments on standard benchmarks demonstrate that MOVAI achieves state-of-the-art performance, improving video quality metrics by 15.3% in LPIPS, 12.7% in FVD, and 18.9% in user preference studies compared to existing methods. Our framework shows particular strength in generating complex multi-object scenes with realistic temporal dynamics and fine-grained semantic control.

文本生成视频时序一致性扩散模型场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。