用大模型自动设计音乐混音效果链,让AI懂音乐后期
LLM2Fx-Tools: Tool Calling For Music Post-Production
- 大模型通过思维链规划,自动选择音效类型、顺序和参数
- 能从原始与处理音频对中推断出完整效果链,准确率显著提升
- 适合音乐制作人、AI音频工具开发者快速生成专业混音方案
本文提出 LLM2Fx-Tools,一个基于大语言模型的多模态工具调用框架,可为音乐后期制作生成可执行的音频效果链(Fx-chain)。该框架利用大语言模型理解音频输入,结合思维链(CoT)规划,自动选择音效类型、确定顺序并估计参数。研究还构建了 LP-Fx 数据集,包含结构化 CoT 注释和音效模块工具调用的指令遵循数据。实验表明,通过自回归序列建模、工具调用与 CoT 推理,系统能从原始与处理音频对中推断出完整的效果链。在风格迁移场景下,系统可将参考源的音效信息迁移到新内容。此外,基于大模型作为裁判的评估显示,本方法能生成合理且适切的音乐制作推理与响应。据我们所知,这是首个将大模型工具调用应用于音效模块的工作,实现了可解释、可控的音乐制作。
原文摘要 · Abstract (English)
This paper introduces LLM2Fx-Tools, a multimodal tool-calling framework that generates executable sequences of audio effects (Fx-chain) for music post-production. LLM2Fx-Tools uses a large language model (LLM) to understand audio inputs, select audio effects types, determine their order, and estimate parameters, guided by chain-of-thought (CoT) planning. We also present LP-Fx, a new instruction-following dataset with structured CoT annotations and tool calls for audio effects modules. Experiments show that LLM2Fx-Tools can infer an Fx-chain and its parameters from pairs of unprocessed and processed audio, enabled by autoregressive sequence modeling, tool calling, and CoT reasoning. We further validate the system in a style transfer setting, where audio effects information is transferred from a reference source and applied to new content. Finally, LLM-as-a-judge evaluation demonstrates that our approach generates appropriate CoT reasoning and responses for music production queries. To our knowledge, this is the first work to apply LLM-based tool calling to audio effects modules, enabling interpretable and controllable music production.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。