arXiv:2607.21857cs.SD2026-07

用智能体构建可控制的声音场景,实现精准音频生成与大规模音语数据构建。

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

论文配图:SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision
图 1 · 摘自论文原文
  • 通过大模型智能体分步规划声音场景,明确分解生成流程
  • 生成音频质量媲美基线,且训练数据提升下游推理性能
  • 支持人机协作编辑,适合需要可控音频生成的研究者

我们提出一种智能体驱动的声音场景构建框架,用于可控的组合式音频生成。该框架显式分解了场景规划、声源选择、时间布局和渲染等传统单次文本到音频模型中隐含的步骤。基于大语言模型的智能体将用户意图转化为可执行的场景计划,通过检索和按需生成获取音源资产,渲染出可控制的多事件混合音频,并导出对齐的场景元数据。框架还支持人机协同交互,允许用户引导工具选择和编辑场景计划。这些组件共同提供了一种可检查、可复用的可控声音场景合成方法,以及可扩展的音语数据构建方案。听觉实验与客观指标显示,其生成效果媲美文本到音频基线;而使用智能体生成数据训练的模型在下游音频推理任务中持续优于仅用真实数据训练的基线。代码、演示及听测结果详见 https://haozhang6720.github.io/SoundscapeAgentDemoPage/。

原文摘要 · Abstract (English)

We present an agentic soundscape construction framework for controllable compositional audio generation that makes explicit the scene planning, source selection, temporal layout, and rendering steps typically handled implicitly by single-shot text-to-audio models. An LLM-based agent converts user intent into an executable scene plan, acquires assets through retrieval and on-demand generation, renders controllable multi-event mixtures, and exports aligned scene metadata. The framework also supports human-in-the-loop interaction through user-guided tool selection and editable scene plans. Together, these components provide an inspectable and reusable approach to controllable soundscape synthesis and scalable audio-language data construction. Listener studies and objective metrics demonstrate competitive generation performance against text-to-audio baselines, while models trained with agent-generated data consistently outperform real-only baselines in downstream audio reasoning. Code, demos, and listening-test results are available at https://haozhang6720.github.io/SoundscapeAgentDemoPage/.

声音生成智能体音语对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。