让对话自动生成多视角分镜,提升影视化叙事质量
Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling
- 用大模型+扩散架构构建无训练框架,分三步生成分镜
- 在多个指标上超越现有方法,分镜连贯性与电影感更强
- 适合影视创作、AI编剧等需要快速可视化对话的场景
近年来,AI驱动的叙事技术在视频生成和故事可视化方面取得进展。然而,将以对话为中心的脚本转化为连贯的分镜仍面临挑战,原因包括脚本信息不足、物理上下文理解有限以及电影原则整合困难。为此,我们提出对话可视化新任务,将对话脚本转化为动态多视角分镜。我们引入无需训练的多模态框架 Dialogue Director,包含剧本导演、摄影师和分镜生成器三个模块。该框架利用大模型与基于扩散的架构,结合思维链推理、检索增强生成和多视角合成等技术,提升对脚本的理解、物理上下文的认知及电影知识的融合能力。实验表明,Dialogue Director 在脚本解析、物理世界理解与电影原则应用方面均优于当前最优方法,显著提升了对话驱动故事可视化的质量和可控性。
原文摘要 · Abstract (English)
Recent advances in AI-driven storytelling have enhanced video generation and story visualization. However, translating dialogue-centric scripts into coherent storyboards remains a significant challenge due to limited script detail, inadequate physical context understanding, and the complexity of integrating cinematic principles. To address these challenges, we propose Dialogue Visualization, a novel task that transforms dialogue scripts into dynamic, multi-view storyboards. We introduce Dialogue Director, a training-free multimodal framework comprising a Script Director, Cinematographer, and Storyboard Maker. This framework leverages large multimodal models and diffusion-based architectures, employing techniques such as Chain-of-Thought reasoning, Retrieval-Augmented Generation, and multi-view synthesis to improve script understanding, physical context comprehension, and cinematic knowledge integration. Experimental results demonstrate that Dialogue Director outperforms state-of-the-art methods in script interpretation, physical world understanding, and cinematic principle application, significantly advancing the quality and controllability of dialogue-based story visualization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。