arXiv:2606.14732cs.CVcs.AI2026-06

解决长视频生成中背景稳定与动态流畅的矛盾,让自然画面持久又生动。

Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion

论文配图:Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion
图 1 · 摘自论文原文
  • 用视觉锚点和运动记忆同步保持场景稳定与动态真实
  • 多分钟生成下背景一致性提升,水流火焰等动态更自然连续
  • 适合关注静态场景自然流动生成的研究者与应用开发者

自回归视频扩散模型在长时间生成中常出现画面漂移:静态场景布局随时间偏移,而增强空间稳定的机制又会抑制运动,导致水、火、烟等自然流动停滞。本文研究固定相机下长时序自然视频生成中的稳定性-运动权衡问题,提出Steady-Forcing框架,结合持久视觉锚点(V-Sink)、指数移动平均运动记忆(EMA-Sink)、块相对时序编码、周期性缓存净化,以及从Wan2.1-14B教师模型蒸馏而来、带有运动奖励先验的任务聚焦配置。该框架在多分钟自回归生成中同时保持背景身份一致性和视觉上合理的流体动力学表现。七种基线对比显示,Steady-Forcing显著提升背景一致性与成像质量;盲测用户研究也表明其感知稳定性与运动连贯性更强。基准评估还揭示,通用VBench得分对固定相机伪影惩罚不足,且将漂移诱导的光流误判为动态度,未直接惩罚纹理硬化或流体停滞,提示未来需建立面向静态相机自然流动的专用评测标准。

原文摘要 · Abstract (English)

Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene layouts drift, while mechanisms that improve spatial stability tend to suppress motion, causing natural flows such as water, fire, or smoke to stagnate. We study this stability-motion trade-off in fixed-camera long-horizon nature video generation, where the two failure modes can be more clearly separated than in moving-camera settings. We propose Steady-Forcing, a memory and training framework combining a persistent visual anchor (V-Sink), an exponential moving-average motion memory (EMA-Sink), block-relative temporal encoding, periodic cache purification, and distillation from a Wan2.1-14B teacher with motion-rewarded priors under task-focused configurations. Together, these components are designed to preserve background identity while sustaining visually plausible fluid dynamics over multi-minute autoregressive rollouts. Evaluations across seven baselines show that Steady-Forcing improves long horizon background consistency and imaging quality, while a blind user study indicates stronger perceived stability and motion continuity. The benchmark evaluation further suggest that generic VBench aggregate scores under-penalize fixed-camera artifacts as well as rewarding drift-induced optical flow as Dynamic Degree while not directly penalizing texture hardening or flow stagnation - motivating future task-specific benchmarks for static-camera nature-flow evaluation. Project page: https://minar09.github.io/steadyforcing/

视频生成扩散模型长视频自然模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。