将长视频分块布局在网格上,提升多镜头连贯性与生成效率。
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

- 用空间网格拆分长视频,分块并行建模减少时间轴负担。
- 生成6.05倍更多镜头,跨镜头一致性达0.5914,优于现有方法。
- 适合需要长视频多镜头连贯生成的研究与应用者。
生成长时序多镜头视频需保证单镜头内运动连贯且镜头间叙事一致。现有视频生成模型倾向于连续运动,在单一时间轴上打包完整叙事时难以呈现完整镜头集合。本文提出MovieGrid多网格后训练范式,将长视频分解为较短的时序片段,并在空间网格中排列进行联合建模。该设计降低每条时间轴处理的镜头数,同时支持片段间的全局信息交互。我们基于1,000个长视频构建了多网格长视频(MGLV)数据集,经源视频采集、层次化分割、网格视频构造与角色感知故事标注,生成54,000组带故事提示的网格视频。采用无噪声随机网格训练,保留部分片段作为干净上下文以去噪其余片段;网格嵌入编码布局结构,角色感知故事提示关联重复实体,网格边界损失稳定布局。在相同标记预算下,MovieGrid在1,616帧视频中生成的镜头数是时序拼接方法的6.05倍。在涵盖五个真实场景的基准上,其单镜头一致性达0.9131(优于HoloCine的0.8086),跨镜头一致性达0.5914(优于StoryMem的0.5384)。通过单次或多次生成,可进一步扩展视频长度且性能损失极小。
原文摘要 · Abstract (English)
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across shots. Existing video generators favor continuous motion and struggle to present complete shot sets when an entire narrative is packed along one temporal axis. We propose MovieGrid, a Multi-Grid Post-Training paradigm that decomposes a long video into shorter, temporally ordered chunks and arranges them on a spatial grid for joint modeling. This design reduces the number of shots handled by each temporal axis while enabling global information exchange across chunks. We construct the Multi-Grid Long Video (MGLV) dataset from 1,000 long-form videos using source video collection, hierarchical segmentation, grid video construction, and character-aware story annotation, producing 54K grid videos paired with story prompts. Our Noise-Free Random-Grid Training retains a random subset of chunks as clean visual context for denoising the remaining chunks. Grid Embedding encodes grid structure, character-aware Story Prompts link recurring entities, and Grid Boundary Loss stabilizes layouts. Under the same token budget, MovieGrid generates 6.05 times more shots than Temporal Packing in a 1,616-frame video. On a benchmark spanning five real-world categories, it achieves state-of-the-art intra-shot consistency (0.9131 versus 0.8086 for HoloCine) and inter-shot consistency (0.5914 versus 0.5384 for StoryMem). MovieGrid can further scale video length with minimal compromise through single or multiple generations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。