arXiv:2604.19679cs.CV2026-04中稿 · ECCV被引 4

让音视频生成同时受图像、声音等多模态控制,实现精准同步。

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

论文配图:MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
图 1 · 摘自论文原文
  • 通过双流注入机制,融合视觉与音频控制信号。
  • 可独立调节视觉与声音影响强度,生成一致身份与音色的音视频。
  • 支持角色身份、声音特质、姿态、场景布局的细粒度控制,适合内容创作者。

近期基于扩散变换器(DiTs)的进展实现了高质量的音视频联合生成,可在单个模型中生成同步音视频。然而,现有可控生成框架通常仅支持视频控制,限制了全面的可控性,并常导致跨模态对齐不佳。为此,我们提出MMControl,实现音视频联合生成中的多模态控制。该方法引入双流条件注入机制,将参考图像、参考音频、深度图、姿态序列等视觉与声学控制信号,通过旁路分支注入到联合音视频扩散变换器中,使模型在结构约束下同时生成身份一致的视频与音色一致的音频。此外,我们设计了模态特定引导缩放机制,可在推理时动态独立调整视觉与声学条件的影响强度。大量实验表明,MMControl在角色身份、语音音色、身体姿态和场景布局方面实现了精细、可组合的控制。

原文摘要 · Abstract (English)

Recent advances in Diffusion Transformers (DiTs) have enabled high-quality joint audio-video generation, producing videos with synchronized audio within a single model. However, existing controllable generation frameworks are typically restricted to video-only control. This restricts comprehensive controllability and often leads to suboptimal cross-modal alignment. To bridge this gap, we present MMControl, which enables users to perform Multi-Modal Control in joint audio-video generation. MMControl introduces a dual-stream conditional injection mechanism. It incorporates both visual and acoustic control signals, including reference images, reference audio, depth maps, and pose sequences, into a joint generation process. These conditions are injected through bypass branches into a joint audio-video Diffusion Transformer, enabling the model to simultaneously generate identity-consistent video and timbre-consistent audio under structural constraints. Furthermore, we introduce modality-specific guidance scaling, which allows users to independently and dynamically adjust the influence strength of each visual and acoustic condition at inference time. Extensive experiments demonstrate that MMControl achieves fine-grained, composable control over character identity, voice timbre, body pose, and scene layout in joint audio-video generation.

音视频生成扩散模型多模态控制条件生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。