arXiv:2503.07217cs.SDcs.CV2025-03被引 2

用多智能体协作生成电影级音效,区分现场与背景声音。

ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation

  • 设计声导、拟音师等智能体,分角色协同生成音效。
  • 通过预测音量、音高、音色控制信号,实现音频与画面同步。
  • 适合影视音效自动化、跨模态生成研究者使用。

现有文本或视频驱动的音频生成方法主要关注模态对齐,但难以处理包含多个场景的电影叙事任务——其中‘在场’声音需与画面时间对齐,而‘离场’声音则需配合环境音与背景音乐。受专业电影制作启发,本文提出多智能体框架ReelWave,由自主声导代理主导,与其他智能体进行多轮对话完成音效生成。针对在场声音,系统检测视频中说话人物后,训练预测模型输出可解释的时间变化音频控制信号(响度、音高、音色),由拟音师代理用于条件化交叉注意力模块生成音效;拟音师协同作曲家与配音师代理,自动生成离场声音以完善整体作品。为实现音频语言模型的时间对齐,在ReelWave中将文本/视频条件分解为原子级、具体的声音生成指令,并与视觉同步。该框架能基于电影片段生成丰富且相关的音频内容。

原文摘要 · Abstract (English)

Current audio generation conditioned by text or video focuses on aligning audio with text/video modalities. Despite excellent alignment results, these multimodal frameworks still cannot be directly applied to compelling movie storytelling involving multiple scenes, where "on-screen" sounds require temporally-aligned audio generation, while "off-screen" sounds contribute to appropriate environment sounds accompanied by background music when applicable. Inspired by professional movie production, this paper proposes a multi-agentic framework for audio generation supervised by an autonomous Sound Director agent, engaging multi-turn conversations with other agents for on-screen and off-screen sound generation through multimodal LLM. To address on-screen sound generation, after detecting any talking humans in videos, we capture semantically and temporally synchronized sound by training a prediction model that forecasts interpretable, time-varying audio control signals: loudness, pitch, and timbre, which are used by a Foley Artist agent to condition a cross-attention module in the sound generation. The Foley Artist works cooperatively with the Composer and Voice Actor agents, and together they autonomously generate off-screen sound to complement the overall production. Each agent takes on specific roles similar to those of a movie production team. To temporally ground audio language models, in ReelWave, text/video conditions are decomposed into atomic, specific sound generation instructions synchronized with visuals when applicable. Consequently, our framework can generate rich and relevant audio content conditioned on video clips extracted from movies.

音效生成多智能体电影制作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。