arXiv:2411.14627cs.SDcs.AI2024-11被引 1

用生成式AI降低音乐创作门槛,让普通人也能轻松创作音频内容。

Generative AI for Music and Audio

  • 构建多轨音乐生成、辅助创作工具与跨模态学习三类技术体系。
  • 实现从零开始生成高质量多轨音乐,支持交互式编辑与风格迁移。
  • 适合音乐创作者、教育者及想快速制作音效的开发者使用。

生成式AI正深刻改变人机交互与内容消费方式。未来十年,该技术将重塑音乐、戏剧、影视、游戏、播客及短视频等领域的音频内容创作模式。本文聚焦生成式AI在音乐与音频领域的三大方向:1)多轨音乐生成,2)辅助音乐创作工具,3)音频与音乐的多模态学习。研究旨在回答两个核心问题:1)AI如何帮助专业人士或爱好者创作音乐与音频?2)AI能否以类似人类学习音乐的方式进行创作?长期目标是降低音乐创作门槛,推动音频内容创作的普及化。

原文摘要 · Abstract (English)

Generative AI has been transforming the way we interact with technology and consume content. In the next decade, AI technology will reshape how we create audio content in various media, including music, theater, films, games, podcasts, and short videos. In this dissertation, I introduce the three main directions of my research centered around generative AI for music and audio: 1) multitrack music generation, 2) assistive music creation tools, and 3) multimodal learning for audio and music. Through my research, I aim to answer the following two fundamental questions: 1) How can AI help professionals or amateurs create music and audio content? 2) Can AI learn to create music in a way similar to how humans learn music? My long-term goal is to lower the barrier of entry for music composition and democratize audio content creation

音乐生成生成式AI音频创作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。