arXiv:2505.08350cs.CVcs.AI2025-05被引 9

用双向生成框架让长篇故事画面连贯可编辑

STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives

  • 双向上下文融合,确保角色和情节连续性
  • 生成画面一致性、叙事丰富度优于开源模型
  • 支持人工修改与长序列扩展,适合创作类应用

本文提出StoryAnchors,一种统一框架,用于生成高质量、多场景且具有强时序一致性的故事画面。该框架采用双向故事生成器,融合过去与未来上下文,确保叙事过程中角色延续性、时序连贯性和场景过渡自然。通过引入特定条件区分故事画面生成与标准视频合成,提升场景多样性与叙事深度。为进一步优化生成质量,整合了多事件故事画面标注与渐进式训练策略,使模型能捕捉整体叙事脉络与事件级动态。该方法支持可编辑、可扩展的故事画面生成,便于人工干预并生成更长更复杂的序列。大量实验表明,StoryAnchors在一致性、叙事连贯性和场景多样性方面超越现有开源模型,其叙事一致性与故事丰富度媲美GPT-4o。最终,StoryAnchors推动了以故事驱动的画面生成边界,为未来研究提供了可扩展、灵活且高度可编辑的基础。

原文摘要 · Abstract (English)

This paper introduces StoryAnchors, a unified framework for generating high-quality, multi-scene story frames with strong temporal consistency. The framework employs a bidirectional story generator that integrates both past and future contexts to ensure temporal consistency, character continuity, and smooth scene transitions throughout the narrative. Specific conditions are introduced to distinguish story frame generation from standard video synthesis, facilitating greater scene diversity and enhancing narrative richness. To further improve generation quality, StoryAnchors integrates Multi-Event Story Frame Labeling and Progressive Story Frame Training, enabling the model to capture both overarching narrative flow and event-level dynamics. This approach supports the creation of editable and expandable story frames, allowing for manual modifications and the generation of longer, more complex sequences. Extensive experiments show that StoryAnchors outperforms existing open-source models in key areas such as consistency, narrative coherence, and scene diversity. Its performance in narrative consistency and story richness is also on par with GPT-4o. Ultimately, StoryAnchors pushes the boundaries of story-driven frame generation, offering a scalable, flexible, and highly editable foundation for future research.

故事生成视频合成时序一致可编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。