让用户通过对话逐步优化音乐,而非反复重生成。
MusiChat: Vibe Composing for Music Creation

- 分层生成架构分离结构与表现,支持风格变换和结构保留编辑。
- 对话中准确理解单轮和多轮指令,正确率达95.31%与100%。
- 适合音乐创作人、作曲新手,提升人机协作创编体验。
AI音乐生成虽能通过自然语言生成完整乐曲,但多数系统采用提示-重生成模式,难以实现迭代优化。本文提出MusiChat,一种基于对话的氛围音乐共创系统,通过自然语言交互实现人机协同创作与渐进式修改。其核心是分层可控音乐生成框架,将歌词对齐的结构生成与表达性细节实现分离,支持灵活风格转换与结构保持编辑。系统通过记忆增强架构整合大语言模型与混合符号音乐引擎,持续维护当前创作状态与用户历史。混合意图路由机制高效解析精确修改与开放创意请求。与从头重生成不同,MusiChat在保留音乐结构与用户意图的前提下增量演化作品。客观分析与人类评估显示,单轮与多轮交互准确率分别达95.31%与100%,旋律自然度与音乐质量的喜好比分别为2:1与3:1。结果表明,MusiChat有效支持连贯多轮音乐创作与对话式人机共创。
原文摘要 · Abstract (English)
Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most existing systems follow a prompt-and-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas. We present MusiChat, a conversational vibe composing system that enables collaborative human-AI music creation through natural-language interaction and iterative refinement. At the core of MusiChat is a hierarchical controllable music generation framework that separates lyric-aligned musical structure generation from expressive surface realization, allowing flexible stylistic transformations and structure-preserving edits. The system integrates a large language model with a hybrid symbolic music engine through a memory-augmented architecture that maintains the active composition state and user history across interactions. A hybrid intent-routing mechanism further enables efficient interpretation of both precise musical edits and open-ended creative requests. Rather than regenerating compositions from scratch, MusiChat incrementally transforms an evolving musical artifact while preserving relevant musical structure and user intent. We evaluate MusiChat through objective analysis and human studies, achieving 95.31% and 100% accuracy for single- and multi-turn interactions, respectively, and obtaining like-to-dislike ratios of 2:1 for melody naturalness and 3:1 for musical quality. Our results demonstrate that MusiChat supports coherent multi-turn music authoring and interactive human-AI co-creation through a conversational interface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。