用记忆管理解决第一人称视频生成中的长期一致性难题
EgoLCD: Egocentric Video Generation with Long Context Diffusion
- 采用长短时记忆结合的架构,稳定保持全局上下文
- 在EgoVid-5M上实现最优感知质量和时间一致性
- 适合构建具身智能的可扩展世界模型研究者
生成长时、连贯的第一人称视频极具挑战,因手物交互和流程性任务需可靠长期记忆。现有自回归模型存在内容漂移问题,导致物体身份与场景语义随时间退化。为此,我们提出EgoLCD,一种端到端的第一人称长上下文视频生成框架,将长视频合成视为高效稳定的记忆管理问题。EgoLCD结合长时稀疏键值缓存以维持全局上下文,辅以基于注意力的短时记忆,并通过LoRA实现局部适应。记忆调节损失强制一致的记忆使用,结构化叙事提示提供显式时间引导。在EgoVid-5M基准上的大量实验表明,EgoLCD在感知质量与时间一致性方面均达到当前最佳表现,有效缓解生成遗忘,为构建具身智能的可扩展世界模型迈出关键一步。
原文摘要 · Abstract (English)
Generating long, coherent egocentric videos is difficult, as hand-object interactions and procedural tasks require reliable long-term memory. Existing autoregressive models suffer from content drift, where object identity and scene semantics degrade over time. To address this challenge, we introduce EgoLCD, an end-to-end framework for egocentric long-context video generation that treats long video synthesis as a problem of efficient and stable memory management. EgoLCD combines a Long-Term Sparse KV Cache for stable global context with an attention-based short-term memory, extended by LoRA for local adaptation. A Memory Regulation Loss enforces consistent memory usage, and Structured Narrative Prompting provides explicit temporal guidance. Extensive experiments on the EgoVid-5M benchmark demonstrate that EgoLCD achieves state-of-the-art performance in both perceptual quality and temporal consistency, effectively mitigating generative forgetting and representing a significant step toward building scalable world models for embodied AI. Code: https://github.com/AIGeeksGroup/EgoLCD. Website: https://aigeeksgroup.github.io/EgoLCD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。