arXiv:2602.11564cs.CV2026-02被引 16

提出双频专家级潜空间级联框架,实现超高清视频生成

LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

  • 分三阶段生成:先低分辨率运动建模,再潜空间上采样,最后融合高低频专家细化细节
  • 在1024×1024分辨率下生成视频,显著提升画质与运动一致性
  • 适合需要高保真视频生成的科研与工业场景

近期视频扩散模型虽大幅提升了视觉质量,但超高清(UHR)视频生成仍面临运动建模、语义规划与细节合成等多重挑战。为此,我们提出LUVE——基于双频专家的潜空间级联超高清视频生成框架。该框架采用三阶段架构:首先生成低分辨率运动一致的潜在表示,随后在潜空间直接进行分辨率上采样以降低内存与计算开销,最后通过低频与高频专家协同增强语义连贯性与精细细节。大量实验表明,LUVE在超高清视频生成中实现了更优的逼真度与内容保真度,消融实验进一步验证了各模块的有效性。

原文摘要 · Abstract (English)

Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a formidable challenge due to the compounded difficulties of motion modeling, semantic planning, and detail synthesis. To address these limitations, we propose \textbf{LUVE}, a \textbf{L}atent-cascaded \textbf{U}HR \textbf{V}ideo generation framework built upon dual frequency \textbf{E}xperts. LUVE employs a three-stage architecture comprising low-resolution motion generation for motion-consistent latent synthesis, video latent upsampling that performs resolution upsampling directly in the latent space to mitigate memory and computational overhead, and high-resolution content refinement that integrates low-frequency and high-frequency experts to jointly enhance semantic coherence and fine-grained detail generation. Extensive experiments demonstrate that our LUVE achieves superior photorealism and content fidelity in UHR video generation, and comprehensive ablation studies further validate the effectiveness of each component. The project is available at \href{https://unicornanrocinu.github.io/LUVE_web/}{https://github.io/LUVE/}.

视频生成扩散模型超高清潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。