用音频生成逼真人物表演视频,无需文本或图像输入。
Seeing Voices: Generating A-Roll Video from Audio with Mirage
- 基于自注意力机制的统一训练方法,从音频直接生成视频。
- 在说话人视频生成任务中主观质量优于专用模型。
- 适合影视创作、虚拟主播等需要音画同步的应用场景。
从专业影视到用户生成内容,视频的感染力取决于声音与画面的和谐统一。现有视频生成方法要么忽略音频,专注于通用图像序列生成,要么虽兼顾视听但局限于特定场景(如配音重制)。我们提出Mirage,一个端到端的音频到视频基础模型,能仅凭音频输入生成真实、富有表现力的视频内容。当与语音合成(文本转语音,TTS)结合时,可生成连贯的多模态视频。在包含人物讲话的音视频数据上训练,并以含语音的音频为条件,Mirage能生成可信的人物表演视频。核心技术是统一的自注意力模型训练方法,支持从零训练或基于已有权重微调。该方法保持了音频到视频生成的通用性,同时在主观评价上优于采用音频特化架构或损失函数的模型。建议读者通过论文附录和评论区链接亲自体验效果。
原文摘要 · Abstract (English)
From professional filmmaking to user-generated content, creators and consumers have long recognized that the power of video depends on the harmonious integration of what we hear (the video's audio track) with what we see (the video's image sequence). Current approaches to video generation either ignore sound to focus on general-purpose but silent image sequence generation or address both visual and audio elements but focus on restricted application domains such as re-dubbing. We introduce Mirage, an audio-to-video foundation model that excels at generating realistic, expressive output imagery from scratch given an audio input. When integrated with existing methods for speech synthesis (text-to-speech, or TTS), Mirage results in compelling multimodal video. When trained on audio-video footage of people talking (A-roll) and conditioned on audio containing speech, Mirage generates video of people delivering a believable interpretation of the performance implicit in input audio. Our central technical contribution is a unified method for training self-attention-based audio-to-video generation models, either from scratch or given existing weights. This methodology allows Mirage to retain generality as an approach to audio-to-video generation while producing outputs of superior subjective quality to methods that incorporate audio-specific architectures or loss components specific to people, speech, or details of how images or audio are captured. We encourage readers to watch and listen to the results of Mirage for themselves (see paper and comments for links).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。