arXiv:2605.23610cs.CVcs.AI2026-05被引 2

用实体为中心的内存实现多镜头视频生成的高效一致性

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation

论文配图:EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
图 1 · 摘自论文原文
  • 构建实体索引的潜在块记忆库,隔离持久性与临时信息
  • 稀疏标记条件降低计算开销,支持预训练模型直接使用
  • 噪声注入控制外观细节,适合需要精准角色一致性的场景

多镜头视频生成需在保持角色跨镜头外观一致的同时忠实于每段文本提示。现有自回归方法复用已生成帧作为记忆,但全图存储会混淆持久实体信息与临时场景上下文,导致无关信息泄露和高计算成本。本文提出一种以实体为中心的记忆机制,采用实体索引的潜在块存储库;引入与预训练模型兼容的稀疏标记条件,限制自注意力仅作用于与实体相关的标记,降低计算开销。为此设计结构化多镜头脚本格式,并提出预算约束的记忆更新策略,维持紧凑且动态演化的记忆体。此外,为实体表示引入噪声注入机制,实现精细外观控制,防止无关信息泄漏。实验表明,该方法在提升提示遵循度和生成效率的同时,有效保持主体一致性。

原文摘要 · Abstract (English)

Multi-shot video generation requires maintaining a consistent appearance of recurring entities across shots while remaining faithful to shot-specific text prompts. Recent autoregressive methods reuse previously generated frames as memory. However, full-frame storage entangles persistent entity information with transient scene context, leading to irrelevant information leakage and high computational cost. We propose an entity-centric memory in the form of an entity-indexed bank of latent patches. We introduce sparse token conditioning compatible with pretrained models, restricting self-attention to entity-relevant tokens and reducing computational cost. To support this, we introduce a structured multi-shot script format. We additionally propose a budgeted memory update strategy to maintain a compact, evolving memory. Finally, we equip the entity representation with a noise-injection mechanism that enables fine-grained appearance control, preventing leakage of irrelevant information. Our method improves prompt adherence and efficiency while preserving subject consistency.

视频生成记忆机制一致性控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。