arXiv:2501.18940cs.CV2025-01被引 1

让视频角色按指定主题实时对话,保持情绪和动作一致。

TV-Dialogue: Crafting Theme-Aware Video Dialogues with Immersive Interaction

  • 构建多模态代理框架,实现角色间沉浸式互动生成对话。
  • 零样本生成任意长度视频的任意主题对话,无需训练。
  • 自建评估基准,验证对话在主题与视觉上的一致性优势。

大语言模型的发展推动了文本与图像对话生成的进步,但基于视频的对话生成仍鲜有探索,面临独特挑战。本文提出主题感知视频对话创作(TVDC)任务,旨在生成与视频内容一致且符合用户指定主题的新对话。我们设计了TV-Dialogue多模态代理框架,通过视频角色间的实时沉浸式互动,确保对话既贴合主题,又与视频中人物的情绪和行为保持一致,从而精准理解视频内容并生成契合主题的新对话。为评估生成效果,我们构建了一个多粒度评估基准,具备高准确率、可解释性与可靠性,实验证明其在自建数据集上优于直接使用现有LLM。大量实验表明,TV-Dialogue可在零样本条件下,为任意长度和主题的视频生成对话,无需训练。研究结果凸显其在视频重制、电影配音及下游多模态任务中的应用潜力。

原文摘要 · Abstract (English)

Recent advancements in LLMs have accelerated the development of dialogue generation across text and images, yet video-based dialogue generation remains underexplored and presents unique challenges. In this paper, we introduce Theme-aware Video Dialogue Crafting (TVDC), a novel task aimed at generating new dialogues that align with video content and adhere to user-specified themes. We propose TV-Dialogue, a novel multi-modal agent framework that ensures both theme alignment (i.e., the dialogue revolves around the theme) and visual consistency (i.e., the dialogue matches the emotions and behaviors of characters in the video) by enabling real-time immersive interactions among video characters, thereby accurately understanding the video content and generating new dialogue that aligns with the given themes. To assess the generated dialogues, we present a multi-granularity evaluation benchmark with high accuracy, interpretability and reliability, demonstrating the effectiveness of TV-Dialogue on self-collected dataset over directly using existing LLMs. Extensive experiments reveal that TV-Dialogue can generate dialogues for videos of any length and any theme in a zero-shot manner without training. Our findings underscore the potential of TV-Dialogue for various applications, such as video re-creation, film dubbing and its use in downstream multimodal tasks.

视频对话多模态主题生成零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。