用记忆机制让视频模型连贯生成多镜头长故事
StoryMem: Multi-shot Long Video Storytelling with Memory
- 用动态记忆库存关键帧,指导多镜头视频生成
- 生成视频跨镜头一致性显著提升,保持高画质与提示符合度
- 适合需要连贯长视频创作的场景,如影视预演
视觉讲故事需要生成具有电影级质量且长程一致性的多镜头视频。受人类记忆启发,我们提出 StoryMem,将长视频故事生成重构为基于显式视觉记忆的迭代镜头合成过程,使预训练的单镜头视频扩散模型具备多镜头叙事能力。该方法通过新颖的 Memory-to-Video (M2V) 设计,维护一个紧凑且动态更新的历史生成镜头关键帧记忆库。记忆内容通过潜空间拼接和负 RoPE 偏移注入单镜头视频扩散模型,仅需 LoRA 微调即可实现。结合语义关键帧选择策略与审美偏好过滤,确保记忆信息丰富且稳定。此外,该框架天然支持平滑镜头过渡与定制化故事生成。为便于评估,我们构建了 ST-Bench,一个多样化的多镜头视频讲故事基准。大量实验表明,StoryMem 在跨镜头一致性上优于现有方法,同时保持高美学质量与提示遵循度,标志着迈向连贯分钟级视频叙事的重要一步。
原文摘要 · Abstract (English)
Visual storytelling requires generating multi-shot videos with cinematic quality and long-range consistency. Inspired by human memory, we propose StoryMem, a paradigm that reformulates long-form video storytelling as iterative shot synthesis conditioned on explicit visual memory, transforming pre-trained single-shot video diffusion models into multi-shot storytellers. This is achieved by a novel Memory-to-Video (M2V) design, which maintains a compact and dynamically updated memory bank of keyframes from historical generated shots. The stored memory is then injected into single-shot video diffusion models via latent concatenation and negative RoPE shifts with only LoRA fine-tuning. A semantic keyframe selection strategy, together with aesthetic preference filtering, further ensures informative and stable memory throughout generation. Moreover, the proposed framework naturally accommodates smooth shot transitions and customized story generation applications. To facilitate evaluation, we introduce ST-Bench, a diverse benchmark for multi-shot video storytelling. Extensive experiments demonstrate that StoryMem achieves superior cross-shot consistency over previous methods while preserving high aesthetic quality and prompt adherence, marking a significant step toward coherent minute-long video storytelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。