arXiv:2601.22160cs.GRcs.AI2026-01

无需训练,通过记忆与匹配实现动作连贯的动画生成。

Screen, Cache, and Match: A Training-Free Causality-Consistent Reference Frame Framework for Human Animation

  • 用动态参考记忆库保存历史帧,减少身份漂移。
  • 跨视频片段对齐去噪轨迹,提升时序一致性。
  • 可无缝接入各类扩散模型,适合长期动画生成研究者。

人体动画旨在生成长时间序列中时间连贯且视觉一致的视频,但建模长程依赖关系同时保持帧质量仍具挑战。受人类利用过往观察理解当前动作能力启发,我们提出FrameCache——一种无需训练、符合因果关系的参考帧框架。该框架通过两种互补机制将历史生成结果转化为因果引导:首先在参考层,采用创新的Screen-Cache-Match(SCM)策略构建动态高质量参考记忆,确保运动一致的外观引导以减少身份漂移;其次在生成层,引入轨迹感知自回归生成(TAAG)机制,通过重叠感知的潜在传播与双域融合策略,对齐相邻视频块间的去噪轨迹,实现低频结构布局与高频纹理细节的无缝融合。在标准基准上的大量实验表明,FrameCache在不依赖特定模型的前提下,持续提升时序连贯性与视觉稳定性,且可无缝集成至多种扩散基线模型。代码将公开。

原文摘要 · Abstract (English)

Human animation aims to generate temporally coherent and visually consistent videos over long sequences, yet modeling long-range dependencies while preserving frame quality remains challenging. Inspired by the human ability to leverage past observations for interpreting ongoing actions, we propose FrameCache, a training-free, causality-consistent reference frame framework. FrameCache explicitly converts historical generation results into causal guidance through two complementary mechanisms. First, at the reference level, a novel Screen-Cache-Match (SCM) strategy constructs a dynamic, high-quality reference memory, ensuring motion-consistent appearance guidance to reduce identity drift. Second, at the generative level, a Trajectory-Aware Autoregressive Generation (TAAG) mechanism aligns denoising trajectories across adjacent video chunks. This is achieved through an overlap-aware latent propagation and a dual-domain fusion strategy that seamlessly blends low-frequency structural layouts with high-frequency textural details. Extensive experiments on standard benchmarks demonstrate that FrameCache consistently improves temporal coherence and visual stability while integrating seamlessly with diverse diffusion baselines. Code will be made publicly available.

人体动画扩散模型时序一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。