arXiv:2504.18448cs.CV2025-04ICCV被引 4

通过分解与协作噪声,提升多视角视频生成的一致性。

NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and Collaboration

  • 将噪声分层分解为场景级与个体级共享/残差成分,捕捉运动特性。
  • 引入跨视图与跨帧协作机制,增强时空一致性与视频质量。
  • 在多个公开数据集上表现领先,适合高一致性要求的视频生成任务。

高质量视频生成在影视制作和自动驾驶等领域至关重要,但保持时空一致性仍是难题。现有方法多依赖注意力机制或修改噪声,忽视了全局时空信息对一致性的潜在作用。本文提出NoiseController,包含多层级噪声分解、多帧噪声协作与联合去噪三部分。首先将初始噪声分解为场景级前景/背景噪声,捕捉不同运动特性;进一步将每类场景噪声分解为个体级共享与残差成分,共享部分维持一致性,残差部分保留多样性。在多帧噪声协作中,设计跨视图时空协作矩阵与视内影响协作矩阵,以建模跨视图相互作用与历史帧间影响。联合去噪采用两个并行的去噪U-Net,分别处理场景级噪声并互相增强生成效果。在多个公开视频生成数据集及下游任务上的实验表明,该方法达到当前最优性能。

原文摘要 · Abstract (English)

High-quality video generation is crucial for many fields, including the film industry and autonomous driving. However, generating videos with spatiotemporal consistencies remains challenging. Current methods typically utilize attention mechanisms or modify noise to achieve consistent videos, neglecting global spatiotemporal information that could help ensure spatial and temporal consistency during video generation. In this paper, we propose the NoiseController, consisting of Multi-Level Noise Decomposition, Multi-Frame Noise Collaboration, and Joint Denoising, to enhance spatiotemporal consistencies in video generation. In multi-level noise decomposition, we first decompose initial noises into scene-level foreground/background noises, capturing distinct motion properties to model multi-view foreground/background variations. Furthermore, each scene-level noise is further decomposed into individual-level shared and residual components. The shared noise preserves consistency, while the residual component maintains diversity. In multi-frame noise collaboration, we introduce an inter-view spatiotemporal collaboration matrix and an intra-view impact collaboration matrix , which captures mutual cross-view effects and historical cross-frame impacts to enhance video quality. The joint denoising contains two parallel denoising U-Nets to remove each scene-level noise, mutually enhancing video generation. We evaluate our NoiseController on public datasets focusing on video generation and downstream tasks, demonstrating its state-of-the-art performance.

视频生成噪声分解时空一致性多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。