用混合记忆机制提升长视频生成的连贯性与互动性
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
- 结合状态空间模型与局部窗口,实现全局动态记忆与细节保留
- 在分钟级长视频生成中保持高运动稳定性与内容多样性
- 支持交互式提示控制,线性扩展于视频长度,适合长视频应用
自回归扩散模型通过逐帧生成实现流式、交互式长视频生成,但长时间跨度下仍面临误差累积、运动漂移和内容重复等挑战。本文从记忆视角出发,将视频合成视为需协调短时与长时上下文的递归动力过程。提出VideoSSM,一种融合自回归扩散与混合状态空间记忆的长视频生成模型:状态空间模型(SSM)作为全序列演化中的全局场景动态记忆,而上下文窗口则提供局部记忆以捕捉运动线索与精细细节。该混合设计在不产生固定重复模式的前提下维持全局一致性,支持提示自适应交互,并实现与序列长度线性增长的计算开销。在短时与长时基准测试中,VideoSSM在自回归视频生成器中展现出领先的时空一致性与运动稳定性,尤其在分钟级生成任务中表现优异,实现了内容多样性与交互式提示控制,建立了一个可扩展、具记忆感知能力的长视频生成框架。
原文摘要 · Abstract (English)
Autoregressive (AR) diffusion enables streaming, interactive long-video generation by producing frames causally, yet maintaining coherence over minute-scale horizons remains challenging due to accumulated errors, motion drift, and content repetition. We approach this problem from a memory perspective, treating video synthesis as a recurrent dynamical process that requires coordinated short- and long-term context. We propose VideoSSM, a Long Video Model that unifies AR diffusion with a hybrid state-space memory. The state-space model (SSM) serves as an evolving global memory of scene dynamics across the entire sequence, while a context window provides local memory for motion cues and fine details. This hybrid design preserves global consistency without frozen, repetitive patterns, supports prompt-adaptive interaction, and scales in linear time with sequence length. Experiments on short- and long-range benchmarks demonstrate state-of-the-art temporal consistency and motion stability among autoregressive video generator especially at minute-scale horizons, enabling content diversity and interactive prompt-based control, thereby establishing a scalable, memory-aware framework for long video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。