让视频生成记住主角,避免角色变形或消失。
Memento: Reconstruct to Remember for Consistent Long Video Generation

- 用记忆重建来显式保持角色一致性,训练生成与记忆回放同步。
- 在长视频中角色一致性提升显著,跨镜头连贯性更强。
- 适合需要长期角色一致性的视频生成任务,如影视创作。
长视频生成要求重复出现的角色在不同镜头、视角、动作和场景切换中保持一致。现有时间分解方法通过逐镜头生成提升可扩展性,但主要关注下一镜头的合理性,未验证历史记忆是否保留关键身份信息。因此,生成过程中角色特征可能逐渐弱化、覆盖或遗忘。本文提出 Memento,一种基于主体重建引导的框架,将主体保持视为显式身份锚定问题:若记忆库能忠实保存主体,则应仅凭记忆即可重建该主体。Memento 联合训练自回归下一镜头生成与基于记忆的主体重建,利用历史记忆和全局故事描述恢复目标外观。为分离长程主体证据与短程线索,引入双查询记忆机制——一个查询检索身份相关记忆,另一个选择短上下文关键帧以保证连贯性。此外,主体感知的电影级数据流水线通过一致且无代词的主体描述提供精确重建监督。实验表明,Memento 在长期主体一致性、跨镜头连贯性和视觉质量方面均达到当前最佳性能。
原文摘要 · Abstract (English)
Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating videos shot by shot. However, they mainly focus on optimizing plausible next-shot continuations without verifying whether the historical memory preserves identity-critical subject evidence. Consequently, as generation proceeds, recurring subjects may be diluted, overwritten, or forgotten. In this paper, we propose Memento, a subject-reconstruction-guided framework that treats subject preservation as an explicit identity grounding problem, based on the premise that a memory bank faithfully preserving a subject should support reconstructing that subject from memory alone. Specifically, Memento jointly trains autoregressive next-shot generation with memory-based subject reconstruction, recovering target appearances using historical memory and global story captions. To disentangle long-range subject evidence from short-range cues, Memento introduces a dual-query memory mechanism, where one query retrieves identity-relevant memory and the other selects short-context keyframes for coherent continuation. Additionally, a subject-aware cinematic data pipeline provides precise reconstruction supervision via consistent, pronoun-free subject descriptions. Experiments demonstrate that Memento achieves state-of-the-art performance in long-term subject consistency, cross-shot coherence, and visual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。