arXiv:2603.03646cs.CV2026-03被引 4

让视频故事无限续写,角色和场景始终一致。

InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions

  • 分阶段生成视频,确保背景与角色关系持续一致。
  • 在多角色出入镜头时实现平滑转场,准确率超82%。
  • 适合长篇剧情视频生成,尤其关注角色连贯性。

生成具有持续视觉叙事的长篇故事视频仍是视频合成中的重大挑战。本文提出一种新框架、数据集与模型,解决三个关键问题:跨镜头的背景一致性、多主体间的自然转场,以及支持长达一小时叙事的可扩展性。所提方法构建了背景一致的生成流程,保持场景视觉连贯性的同时维护角色身份与空间关系。进一步设计了过渡感知的视频生成模块,可处理多个主体进出画面的复杂场景,突破以往仅支持单主体的局限。为此,我们构建了一个包含10,000个多主体转场序列的合成数据集,涵盖未被充分覆盖的动态场景组合。在VBench评测中,InfinityStory取得最高背景一致性(88.94)、最高主体一致性(82.11),整体平均排名最优(2.80),展现出更强的稳定性、更流畅的转场与更好的时间连贯性。

原文摘要 · Abstract (English)

Generating long-form storytelling videos with consistent visual narratives remains a significant challenge in video synthesis. We present a novel framework, dataset, and a model that address three critical limitations: background consistency across shots, seamless multi-subject shot-to-shot transitions, and scalability to hour-long narratives. Our approach introduces a background-consistent generation pipeline that maintains visual coherence across scenes while preserving character identity and spatial relationships. We further propose a transition-aware video synthesis module that generates smooth shot transitions for complex scenarios involving multiple subjects entering or exiting frames, going beyond the single-subject limitations of prior work. To support this, we contribute with a synthetic dataset of 10,000 multi-subject transition sequences covering underrepresented dynamic scene compositions. On VBench, InfinityStory achieves the highest Background Consistency (88.94), highest Subject Consistency (82.11), and the best overall average rank (2.80), showing improved stability, smoother transitions, and better temporal coherence.

视频生成角色一致性长视频转场优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。