arXiv:2511.17323cs.SDcs.AI2025-11中稿 · IEEE Big Data 2025被引 1

用算法驱动的符号音乐核心,从歌词自动生成符合乐理的歌曲。

MusicAIR: A Multimodal AI Music Generation Framework Powered by an Algorithm-Driven Core

  • 基于算法解析歌词与节奏,自动生成完整旋律
  • 生成曲目键位准确率85%,高于人类作曲家的79%
  • 适合音乐创作新手和教育场景使用

生成式AI在音乐生成领域取得显著进展,但多数神经网络模型依赖大规模数据集,引发版权争议与高算力成本问题。为此,我们提出MusicAIR——一种由创新算法驱动的符号音乐核心支撑的多模态音乐生成框架,有效降低版权风险。该框架通过连接歌词与节奏信息,自动推导音乐特征,仅凭歌词即可生成完整、连贯的旋律谱。MusicAIR支持从歌词、文本到图像的多模态音乐生成,产出作品严格遵循乐理、歌词结构与节奏规范。我们开发了名为GenAIM的Web工具,用于实现歌词转歌曲、文本转音乐及图像转音乐功能。实验中,系统生成作品平均键位识别准确率达85%,优于人类作曲家的79%,且与经典作品高度契合,展现出多样且类人化的创作能力。作为协作者工具,GenAIM可作为可靠的作曲助手或教学辅导系统,大幅降低音乐创作门槛,对人工智能音乐生成具有重要创新价值。

原文摘要 · Abstract (English)

Recent advances in generative AI have made music generation a prominent research focus. However, many neural-based models rely on large datasets, raising concerns about copyright infringement and high-performance costs. In contrast, we propose MusicAIR, an innovative multimodal AI music generation framework powered by a novel algorithm-driven symbolic music core, effectively mitigating copyright infringement risks. The music core algorithms connect critical lyrical and rhythmic information to automatically derive musical features, creating a complete, coherent melodic score solely from the lyrics. The MusicAIR framework facilitates music generation from lyrics, text, and images. The generated score adheres to established principles of music theory, lyrical structure, and rhythmic conventions. We developed Generate AI Music (GenAIM), a web tool using MusicAIR for lyric-to-song, text-to-music, and image-to-music generation. In our experiments, we evaluated AI-generated music scores produced by the system using both standard music metrics and innovative analysis that compares these compositions with original works. The system achieves an average key confidence of 85%, outperforming human composers at 79%, and aligns closely with established music theory standards, demonstrating its ability to generate diverse, human-like compositions. As a co-pilot tool, GenAIM can serve as a reliable music composition assistant and a possible educational composition tutor while simultaneously lowering the entry barrier for all aspiring musicians, which is innovative and significantly contributes to AI for music generation.

音乐生成多模态符号音乐算法驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。