arXiv:2605.31603cs.CVcs.AI2026-05

轻量训练+高频渐进衔接,实现高画质视频生成

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models

论文配图:Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
图 1 · 摘自论文原文
  • 训练时用轻量生成器对齐理解模块,降低计算成本
  • 推理时通过渐进式频率桥接,提升视频画质与连贯性
  • 专设基准测试VR-Bench,评估意图到视频的语义一致性

基于连接器的视频统一模型在指令驱动视频生成中表现出强大能力,但将大型高保真生成器整合到统一训练流程中计算开销巨大,限制了视觉质量的提升。为此,我们提出Lumos-Nexus,一种训练高效的统一视频生成框架,能够在显著提升视觉保真度的同时保留强推理驱动生成能力。该框架采用两阶段设计:1)训练阶段仅让轻量生成器与理解模块对齐,学习接受推理驱动的语义控制;2)推理阶段引入统一渐进频率桥接(UPFB),在共享潜在空间中逐步移交生成任务至高容量预训练生成器,实现从粗到细的优化,生成高保真视频而不牺牲推理质量。为填补推理驱动视频生成的评估空白,我们构建了VR-Bench,用于评估模型将推断意图转化为连贯且语义一致视频内容的能力。大量实验表明,Lumos-Nexus在VBench上显著提升视觉真实感与时间连贯性,同时在VR-Bench上展现出强大的基于推理的生成性能。

原文摘要 · Abstract (English)

Connector-based video unified models have demonstrated strong capability in instruction-grounded video synthesis, but integrating a large high-fidelity generator into the unified training loop is computationally prohibitive, limiting achievable visual quality. We therefore propose Lumos-Nexus, a training-efficient unified video generation framework that facilitates the development of strong reasoning-driven generation capabilities while significantly enhancing visual fidelity. Lumos-Nexus adopts a two-stage design: 1) During training, only a lightweight generator is aligned with the understanding block to learn to take in reasoning-driven semantic control. 2) During inference, we introduce Unified Progressive Frequency Bridging (UPFB) to progressively hand off generation to a high-capacity pretrained generator in the shared latent space, enabling coarse-to-fine refinement and producing high-fidelity videos without compromising reasoning quality. To fill the gap in reasoning-driven video generation benchmarks, we introduce VR-Bench, which assesses a model's capability to translate inferred intent into coherent and semantically aligned video content. Extensive experiments demonstrate that Lumos-Nexus achieves substantial gains in visual realism and temporal coherence on VBench, while exhibiting strong reasoning-based generative performance on VR-Bench. Code and models are available at https://jiazheng-xing.github.io/nexus-lumos-home/.

视频生成扩散模型推理生成高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。