arXiv:2505.16819cs.CV2025-05中稿 · the 2026 IEEE Inte…被引 2

让故事角色根据场景提示生成连贯对话,实现有声有情的多模态叙事。

Character-Centered Dialogue Generation from Scene-Level Prompts

  • 用视觉语言模型结合场景提示生成角色台词,确保内容与画面一致。
  • 通过递归叙事库保持角色情感和对话在时间上的连贯性。
  • 无需训练即可适配多种故事场景,适合影视创作与互动叙事应用。

基于场景的视频生成技术已能从结构化提示生成连贯视觉叙事,但角色驱动的对话与语音仍研究不足。本文提出一种模块化框架,将动作级提示转化为视觉与听觉双重支撑的对话,增强场景叙事的表现力。方法每场景输入一对提示:设定描述与角色行为,由Text2Story等模型生成视觉场景后,利用预训练视觉语言编码器提取图像高层语义,结合结构化提示引导大语言模型生成具表现力、角色一致的对白。为维持跨场景的情境与情感一致性,引入递归叙事库——一种面向说话者、时序化的记忆机制,记录角色对话历史。受脚本理论启发,该设计使对话体现角色目标演变、社交关系与叙事身份的变化。最终将每句对白合成个性化语音,生成完整有声多模态视频叙事。整个框架无需训练,可泛化至多样故事场景,提供高效、一致的音视频叙事解决方案。

原文摘要 · Abstract (English)

Recent advances in scene-based video generation enable coherent visual narratives from structured prompts, yet a key aspect of storytelling -- character-driven dialogue and speech -- remains underexplored. We present a modular pipeline that transforms action-level prompts into visually and auditorily grounded dialogue, enriching scene-based storytelling with natural voice and character expression. Our method takes a pair of prompts per scene, defining the setting and character behavior. While a story generation model such as Text2Story produces the visual scene, we focus on generating expressive, character-consistent utterances grounded in both the prompts and a representative scene image. A pretrained vision-language encoder extracts high-level visual semantics, which are combined with structured prompts to guide a large language model for dialogue synthesis. To maintain contextual and emotional consistency across scenes, we introduce a Recursive Narrative Bank, a speaker-aware, temporally structured memory that accumulates each character's dialogue history. Inspired by Script Theory, this design enables dialogue that reflects evolving goals, social context, and narrative roles. Finally, we render each utterance as expressive, character-conditioned speech, producing fully voiced, multimodal video narratives. Our training-free framework generalizes across diverse story settings, providing a scalable solution for coherent, character-grounded audiovisual storytelling.

对话生成多模态叙事角色一致性视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。