arXiv:2506.08003cs.CVcs.AI2025-06NeurIPS被引 14

让视频精准同步音频,分控语音、音效和音乐

Audio-Sync Video Generation with Multi-Stream Temporal Control

  • 拆分音频为语音/音效/音乐三路,分别控制口型、动作和情绪
  • 在六项指标上超越现有方法,视频质量与音画同步性俱佳
  • 适合影视生成、播客可视化等需要精准音画对齐的场景

音频具有固有时序性且与视觉世界紧密关联,是可控视频生成(如电影)的理想控制信号。直接将音频转为视频对理解与可视化丰富音频叙事(如播客或历史录音)也至关重要。然而,现有方法在生成高质量、精确音画同步视频方面表现不足,尤其面对多样复杂的音频类型时。本文提出MTV框架,通过显式分离音频为语音、音效和音乐三路,实现对口型、事件时间与视觉情绪的解耦控制,从而生成细粒度且语义对齐的视频。为支持该框架,我们构建了DEMIX数据集,包含高质量影视视频与分离音频轨道,共五组重叠子集,支持多阶段可扩展训练。大量实验表明,MTV在六项标准指标上达到领先水平,涵盖视频质量、文本-视频一致性及音画对齐能力。

原文摘要 · Abstract (English)

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video is essential for understanding and visualizing rich audio narratives (e.g., Podcasts or historical recordings). However, existing approaches fall short in generating high-quality videos with precise audio-visual synchronization, especially across diverse and complex audio types. In this work, we introduce MTV, a versatile framework for audio-sync video generation. MTV explicitly separates audios into speech, effects, and music tracks, enabling disentangled control over lip motion, event timing, and visual mood, respectively -- resulting in fine-grained and semantically aligned video generation. To support the framework, we additionally present DEMIX, a dataset comprising high-quality cinematic videos and demixed audio tracks. DEMIX is structured into five overlapped subsets, enabling scalable multi-stage training for diverse generation scenarios. Extensive experiments demonstrate that MTV achieves state-of-the-art performance across six standard metrics spanning video quality, text-video consistency, and audio-video alignment. Project page: https://hjzheng.net/projects/MTV/.

音视频生成多模态控制音频分离影视生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。