arXiv:2512.03034cs.CV2025-12International Conf…被引 4

MAViD实现音视频对话的自然生成与理解,支持长时序同步交互。

MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation

  • 分导引-创作双模块架构,分别处理指令分解与响应生成。
  • 融合自回归与扩散模型,实现高质量音视频同步生成。
  • 创新融合模块提升多模态连续片段一致性,适合交互系统开发。

我们提出MAViD,一种用于音视频对话理解与生成的新型多模态框架。现有方法主要聚焦非交互系统,生成受限且不自然的人类语音。该任务的核心挑战在于有效整合理解与生成能力,以及实现无缝的多模态音视频融合。为此,我们设计了导引-创作架构,将对话系统分为两个核心组件:导引模块负责理解、推理并生成包含动作与语音成分的指令,实现对交互的细粒度控制;创作模块则根据指令生成交互式响应。为解决使用双重DiT结构生成长视频时身份、音色和语调不一致的问题,创作模块采用结合自回归(AR)与扩散模型的结构:AR模型负责音频生成,扩散模型保障视频生成质量。此外,我们提出一种新颖的融合模块,增强上下文连续片段与模态间的连接,实现长时间音视频内容的同步生成。大量实验表明,该框架可生成生动且上下文连贯的长时序对话交互,并准确解析用户多模态查询。

原文摘要 · Abstract (English)

We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limited to producing constrained and unnatural human speech. The primary challenge of this task lies in effectively integrating understanding and generation capabilities, as well as achieving seamless multimodal audio-video fusion. To solve these problems, we propose a Conductor-Creator architecture that divides the dialogue system into two primary components. The Conductor is tasked with understanding, reasoning, and generating instructions by breaking them down into motion and speech components, thereby enabling fine-grained control over interactions. The Creator then delivers interactive responses based on these instructions. Furthermore, to address the difficulty of generating long videos with consistent identity, timbre, and tone using dual DiT structures, the Creator adopts a structure that combines autoregressive (AR) and diffusion models. The AR model is responsible for audio generation, while the diffusion model ensures high-quality video generation. Additionally, we propose a novel fusion module to enhance connections between contextually consecutive clips and modalities, enabling synchronized long-duration audio-visual content generation. Extensive experiments demonstrate that our framework can generate vivid and contextually coherent long-duration dialogue interactions and accurately interpret users' multimodal queries.

音视频对话多模态生成扩散模型交互系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。