用隐空间记忆修复动画漂移,实现分钟级人像持续生成
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

- 通过隐空间上下文记忆保持角色身份与动作连贯性
- 10秒生成提升PSNR/SSIM 8%/7%,90秒时达15%/15%
- 仅用轻量LoRA微调,适合长视频生成研究者
我们提出EverAnimate,一种高效的后训练方法,用于生成长时间人类动画视频,同时保持视觉质量与角色身份一致。长时序动画仍具挑战性,因高度动态的人体运动需在相对静态环境中合成,导致分段生成易产生累积漂移:(i) 低层质量漂移,如背景逐渐退化;(ii) 高层语义漂移,如角色身份不一致或视角属性失真。为解决此问题,EverAnimate通过锚定生成过程到持久的隐空间上下文记忆,采用两种互补机制:(i) 持久隐空间传播,在各片段间维持上下文记忆,缓解时间遗忘;(ii) 修复性光流匹配,在采样阶段通过速度调整引入隐式恢复目标,提升片段内保真度。仅需轻量级LoRA微调,EverAnimate在短时(10秒)与长时(90秒)设置下均超越现有最优方法:10秒时PSNR/SSIM提升8%/7%,LPIPS/FID降低22%/11%;90秒时提升增至15%/15%,降低32%/27%。
原文摘要 · Abstract (English)
We propose EverAnimate, an efficient post-training method for long-horizon animated video generation that preserves visual quality and character identity. Long-form animation remains challenging because highly dynamic human motion must be synthesized against relatively static environments, making chunk-based generation prone to accumulated drift: (i) low-level quality drift, such as progressive degradation of static backgrounds, and (ii) high-level semantic drift, such as inconsistent character identity and view-dependent attributes. To address this issue, EverAnimate restores drifted flow trajectories by anchoring generation to a persistent latent context memory, consisting of two complementary mechanisms. (i) Persistent Latent Propagation maintains a context memory across chunks to propagate identity and motion in latent space while mitigating temporal forgetting. (ii) Restorative Flow Matching introduces an implicit restoration objective during sampling through velocity adjustment, improving within-chunk fidelity. With only lightweight LoRA tuning, EverAnimate outperforms state-of-the-art long-animation methods in both short- and long-horizon settings: at 10 seconds, it improves PSNR/SSIM by 8%/7% and reduces LPIPS/FID by 22%/11%; at 90 seconds, the gains increase to 15%/15% and 32%/27%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。