用大模型让音乐生成可自由设定声源位置。
STASE: A spatialized text-to-audio synthesis engine for music generation
- 大模型解析文本空间指令,分离语义理解与物理渲染
- 支持具体坐标和抽象风格两种空间描述方式
- 适合需要精细空间控制的音乐创作与虚拟现实场景
尽管许多文本到音频系统仅生成单声道或固定立体声输出,但实现用户自定义空间属性的音频生成仍具挑战。现有基于深度学习的空间化方法常依赖潜在空间操作,难以直接控制影响空间感知的心理声学参数。为此,我们提出STASE,一个利用大型语言模型(LLM)作为代理来解析文本空间线索的系统。其核心是将语义理解与独立的物理基础空间渲染引擎解耦,实现可解释且用户可控的空间推理。LLM通过两条路径处理提示:(i) 描述性提示,直接映射显式空间信息(如“将主吉他置于45°方位角,10米距离”);(ii) 抽象提示,通过检索增强生成(RAG)模块获取相关空间模板以指导渲染。本文详述了STASE的工作流程,讨论了实现细节,并指出了当前生成空间音频评估中的挑战。
原文摘要 · Abstract (English)
While many text-to-audio systems produce monophonic or fixed-stereo outputs, generating audio with user-defined spatial properties remains a challenge. Existing deep learning-based spatialization methods often rely on latent-space manipulations, which can limit direct control over psychoacoustic parameters critical to spatial perception. To address this, we introduce STASE, a system that leverages a Large Language Model (LLM) as an agent to interpret spatial cues from text. A key feature of STASE is the decoupling of semantic interpretation from a separate, physics-based spatial rendering engine, which facilitates interpretable and user-controllable spatial reasoning. The LLM processes prompts through two main pathways: (i) Description Prompts, for direct mapping of explicit spatial information (e.g., "place the lead guitar at 45° azimuth, 10 m distance"), and (ii) Abstract Prompts, where a Retrieval-Augmented Generation (RAG) module retrieves relevant spatial templates to inform the rendering. This paper details the STASE workflow, discusses implementation considerations, and highlights current challenges in evaluating generative spatial audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。