用自适应记忆建模实现连贯长视频生成,支持文本与图像双重控制。
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- 将多镜头视频生成重构为逐镜头续写任务,引入全局语义记忆
- 在60K数据集上微调后,复杂场景下叙事连贯性领先现有方法
- 适合需要长视频连续生成的创作与交互应用
现实世界中的故事常由多个不连续但语义相连的镜头组成。现有方法因受限于短时窗或单关键帧条件,难以建模长距离跨镜头上下文,导致复杂叙事性能下降。本文提出OneStory,通过将多镜头视频生成重构为下一镜头生成任务,结合预训练图像到视频模型实现强视觉条件化,支持自回归镜头合成。引入帧选择模块构建基于先前镜头信息帧的全局语义记忆,以及自适应调节器进行重要性引导的补丁化处理,生成紧凑上下文直接用于条件输入。我们还构建了一个含参考描述的高质量多镜头数据集,以模拟真实叙事模式,并设计了相应的训练策略。在60K规模数据集上微调预训练图像到视频模型后,OneStory在文本和图像条件设置下均实现了多样且复杂场景下的最佳叙事连贯性,支持可控、沉浸式长视频故事生成。
原文摘要 · Abstract (English)
Storytelling in real-world videos often unfolds through multiple shots -- discontinuous yet semantically connected clips that together convey a coherent narrative. However, existing multi-shot video generation (MSV) methods struggle to effectively model long-range cross-shot context, as they rely on limited temporal windows or single keyframe conditioning, leading to degraded performance under complex narratives. In this work, we propose OneStory, enabling global yet compact cross-shot context modeling for consistent and scalable narrative generation. OneStory reformulates MSV as a next-shot generation task, enabling autoregressive shot synthesis while leveraging pretrained image-to-video (I2V) models for strong visual conditioning. We introduce two key modules: a Frame Selection module that constructs a semantically-relevant global memory based on informative frames from prior shots, and an Adaptive Conditioner that performs importance-guided patchification to generate compact context for direct conditioning. We further curate a high-quality multi-shot dataset with referential captions to mirror real-world storytelling patterns, and design effective training strategies under the next-shot paradigm. Finetuned from a pretrained I2V model on our curated 60K dataset, OneStory achieves state-of-the-art narrative coherence across diverse and complex scenes in both text- and image-conditioned settings, enabling controllable and immersive long-form video storytelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。