arXiv:2512.16954cs.CVcs.AI2025-12被引 2

分阶段生成角色一致的AI视频故事,提升长期连贯性。

Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories

  • 用大语言模型生成剧本,再通过视觉锚点保证角色一致性
  • 去除视觉锚点后角色一致性骤降至0.55,证明其关键作用
  • 发现中西方主题生成存在文化偏见,适合影视创作与伦理研究

当前文本到视频的AI生成在长视频、角色一致性方面仍面临挑战。本文提出一种类电影制作流程的多阶段方法:首先由大语言模型生成详细剧本,指导文本到图像模型为每个角色生成一致的视觉形象,这些图像作为锚点输入视频生成模型,逐场景合成视频。基线对比验证了该分解流程的有效性:移除视觉锚点后,角色一致性评分从7.99暴跌至0.55,证实视觉先验对身份保持至关重要。此外,我们分析了现有模型的文化差异,发现印度与西方主题生成在主体一致性和动态程度上存在显著偏差。

原文摘要 · Abstract (English)

Generating long, cohesive video stories with consistent characters is a significant challenge for current text-to-video AI. We introduce a method that approaches video generation in a filmmaker-like manner. Instead of creating a video in one step, our proposed pipeline first uses a large language model to generate a detailed production script. This script guides a text-to-image model in creating consistent visuals for each character, which then serve as anchors for a video generation model to synthesize each scene individually. Our baseline comparisons validate the necessity of this multi-stage decomposition; specifically, we observe that removing the visual anchoring mechanism results in a catastrophic drop in character consistency scores (from 7.99 to 0.55), confirming that visual priors are essential for identity preservation. Furthermore, we analyze cultural disparities in current models, revealing distinct biases in subject consistency and dynamic degree between Indian vs Western-themed generations.

视频生成角色一致多阶段文化偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。