让文本生成音乐更可控、可编辑,支持对话式迭代与精准属性修改。
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
- 用大模型协调多AI模型,通过对话实现音乐的逐步优化。
- 引入全局属性表,保证修改过程中音乐整体一致性。
- 零样本编辑无需重训,可精准调整风格、情绪、乐器等属性。
AI辅助音乐创作已取得显著进展,但现有系统在迭代创作和精细调整方面仍存挑战。本文提出两套互补方案:首先,设计了Loop Copilot系统,利用大语言模型(LLM)协同多个专用AI模型,通过对话界面实现音乐的交互式生成与优化;其核心是全局属性表,用于记录并维护关键音乐属性,确保各阶段修改保持整体连贯性。其次,提出MusicMagus,一种零样本文本到音乐编辑方法,通过操控预训练扩散模型的隐空间,实现对特定属性(如流派、情绪、乐器)的精确修改,且不改变非目标属性。该方法在保持音乐结构完整性方面表现优异,但在复杂真实音频场景中仍面临挑战。
原文摘要 · Abstract (English)
The field of AI-assisted music creation has made significant strides, yet existing systems often struggle to meet the demands of iterative and nuanced music production. These challenges include providing sufficient control over the generated content and allowing for flexible, precise edits. This thesis tackles these issues by introducing a series of advancements that progressively build upon each other, enhancing the controllability and editability of text-to-music generation models. First, we introduce Loop Copilot, a system that tries to address the need for iterative refinement in music creation. Loop Copilot leverages a large language model (LLM) to coordinate multiple specialised AI models, enabling users to generate and refine music interactively through a conversational interface. Central to this system is the Global Attribute Table, which records and maintains key musical attributes throughout the iterative process, ensuring that modifications at any stage preserve the overall coherence of the music. While Loop Copilot excels in orchestrating the music creation process, it does not directly address the need for detailed edits to the generated content. To overcome this limitation, MusicMagus is presented as a further solution for editing AI-generated music. MusicMagus introduces a zero-shot text-to-music editing approach that allows for the modification of specific musical attributes, such as genre, mood, and instrumentation, without the need for retraining. By manipulating the latent space within pre-trained diffusion models, MusicMagus ensures that these edits are stylistically coherent and that non-targeted attributes remain unchanged. This system is particularly effective in maintaining the structural integrity of the music during edits, but it encounters challenges with more complex and real-world audio scenarios. ...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。