arXiv:2606.07397cs.SD2026-06被引 3

多智能体协作生成复杂音频场景,实现精准可控的音效与音乐合成。

Audio-Oscar: A Multi-Agent System for Complex Audio Scene Generation, Orchestration, and Refinement

论文配图:Audio-Oscar: A Multi-Agent System for Complex Audio Scene Generation, Orchestration, and Refinement
图 1 · 摘自论文原文
  • 分角色、分任务的多智能体协同生成语音、音效、音乐与时间轴。
  • 在新构建的ASG-Bench上生成音频与场景描述匹配度达85.3%。
  • 适合需要高精度音频合成的影视、游戏与虚拟现实开发者。

近年来,语音生成在文本转语音(TTS)、文本转音频(TTA)和文本转音乐(TTM)等任务上取得显著进展。然而,从复杂音频场景描述中生成长时序且可控制的音频仍面临挑战,因这类场景通常需协调语音、音效、音乐、歌曲、时间结构与后期处理。本文提出\textbf{Audio-Oscar},一个用于复杂场景音频生成的多智能体框架。该框架通过一组专业智能体协作,分别负责角色建模与音色设计、语音生成、细粒度时间线规划、模型选择、非语音内容生成及音频后期制作。Audio-Oscar还引入反馈驱动的优化机制。为解决评估基准缺失问题,我们构建了\textbf{ASG-Bench}——一个包含场景描述与参考音频对、以及仅含文本的场景描述的音频场景生成基准。每个场景均标注目标音频事件与时间语句,用于评估生成音频是否忠实还原内容与时间结构。实验表明,Audio-Oscar能有效生成符合复杂场景描述的音频。项目样例见https://audiooscar.github.io/,代码已开源于https://github.com/ziye26/Audio-Oscar。

原文摘要 · Abstract (English)

In recent years, audio generation has made significant progress in tasks such as text-to-speech (TTS), text-to-audio (TTA) and text-to-music (TTM). However, generating long-form and controllable audio from complex audio scene descriptions remains a significant challenge, as such scenes often require coordinated speech, sound effects, music, songs, temporal structure, and post-production. In this work, we introduce \textbf{Audio-Oscar}, a multi-agent framework for generating audio from complex descriptions. Audio-Oscar coordinates a set of specialist agents, each responsible for a different aspect of the audio scene, including character modeling and voice design, speech generation, fine-grained timeline planning, model selection, non-speech generation, and audio post-production. Audio-Oscar further incorporates feedback-driven refinement. In addition, to address the lack of suitable benchmarks for evaluating audio generation from complex audio scene descriptions, we construct \textbf{ASG-Bench}, an Audio Scene Generation Benchmark containing both scene descriptions paired with reference audio and text-only scene descriptions. Each scene is annotated with target audio events and temporal statements to evaluate whether the generated audio faithfully realizes the required scene content and temporal structure. Experimental results show that Audio-Oscar can effectively generate audio that matches complex scene descriptions. Project samples are available at https://audiooscar.github.io/. Our code is available at https://github.com/ziye26/Audio-Oscar.

音频生成多智能体场景合成生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。