arXiv:2503.16400cs.LG2025-03被引 14

通过搜索优质初始噪声,提升视频生成的连贯性与多样性。

ScalingNoise: Scaling Inference-Time Search for Generating Infinite Videos

  • 推理时搜索最优初始噪声,引导扩散过程生成更高质量视频。
  • 单步去噪结合奖励信号,显著提升长视频内容一致性。
  • 适合追求长视频生成质量与多样性的研究者和开发者。

视频扩散模型(VDMs)可生成高质量视频,当前研究多聚焦于训练阶段的规模扩展。然而,推理阶段的扩展关注较少,多数方法仅允许单次生成。近期研究发现存在可提升视频质量的'黄金噪声'。我们发现,引导推理时搜索以识别更优噪声不仅评估当前帧质量,还能通过参考先前多块的锚定帧保留高层物体特征,带来长期价值。分析表明,扩散模型可通过调整去噪步骤灵活调节计算量,即使单步去噪,在奖励信号引导下也具显著长期收益。为此,我们提出ScalingNoise,一种即插即用的推理时搜索策略,通过寻找扩散采样过程中的黄金初始噪声,提升全局内容一致性和视觉多样性。具体地,采用单步去噪将初始噪声转换为片段,并基于先前生成内容的锚定奖励模型评估其长期价值;为保持多样性,从倾斜噪声分布中采样,优先选取有潜力的噪声。实验表明,ScalingNoise显著减少噪声引入误差,实现更连贯、时空一致的视频生成,在多个基准数据集上验证有效。

原文摘要 · Abstract (English)

Video diffusion models (VDMs) facilitate the generation of high-quality videos, with current research predominantly concentrated on scaling efforts during training through improvements in data quality, computational resources, and model complexity. However, inference-time scaling has received less attention, with most approaches restricting models to a single generation attempt. Recent studies have uncovered the existence of "golden noises" that can enhance video quality during generation. Building on this, we find that guiding the scaling inference-time search of VDMs to identify better noise candidates not only evaluates the quality of the frames generated in the current step but also preserves the high-level object features by referencing the anchor frame from previous multi-chunks, thereby delivering long-term value. Our analysis reveals that diffusion models inherently possess flexible adjustments of computation by varying denoising steps, and even a one-step denoising approach, when guided by a reward signal, yields significant long-term benefits. Based on the observation, we proposeScalingNoise, a plug-and-play inference-time search strategy that identifies golden initial noises for the diffusion sampling process to improve global content consistency and visual diversity. Specifically, we perform one-step denoising to convert initial noises into a clip and subsequently evaluate its long-term value, leveraging a reward model anchored by previously generated content. Moreover, to preserve diversity, we sample candidates from a tilted noise distribution that up-weights promising noises. In this way, ScalingNoise significantly reduces noise-induced errors, ensuring more coherent and spatiotemporally consistent video generation. Extensive experiments on benchmark datasets demonstrate that the proposed ScalingNoise effectively improves long video generation.

视频生成扩散模型推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。