开源音乐大模型家族,支持多模态生成与精细控制。
HeartMuLa: A Family of Open Sourced Music Foundation Models
- 构建四组件架构:音频文本对齐、歌词识别、低帧率高保真编码器、基于LLM的音乐生成。
- 7B参数下达到商业级音质,支持自然语言控制不同段落风格。
- 适合音乐生成、视频配乐、多模态创作等研究与应用开发者使用。
我们提出一个开源音乐基础模型家族,旨在推动大规模音乐理解与生成任务的发展。该框架包含四个核心组件:(1) HeartCLAP,用于音频与文本对齐;(2) HeartTranscriptor,针对真实场景优化的鲁棒歌词识别模型;(3) HeartCodec,一种低帧率(12.5 Hz)但高保真的音乐编码器,能捕捉长程音乐结构并保留细微声学细节,支持高效自回归建模;(4) HeartMuLa,基于大语言模型的歌曲生成模型,可在丰富用户可控条件下合成高质量音乐(如文本风格描述、歌词和参考音频)。此外,它提供两种专用模式:(i) 细粒度音乐属性控制,支持用自然语言指定不同乐段(如前奏、主歌、副歌)风格;(ii) 短时吸引人音乐生成,适用于短视频背景音乐。实验表明,模型在扩展至70亿参数后性能显著提升。首次证明仅用学术数据与算力即可复现类似Suno的商业级系统。这些模型有望成为未来研究的重要基准,并促进多模态内容生产落地。
原文摘要 · Abstract (English)
We present a family of open-source Music Foundation Models designed to advance large-scale music understanding and generation across diverse tasks and modalities. Our framework consists of four major components: (1) HeartCLAP, an audio-text alignment model; (2) HeartTranscriptor, a robust lyric recognition model optimized for real-world music scenarios; and (3) HeartCodec, a low-frame-rate (12.5 Hz) yet high-fidelity music codec tokenizer that captures long-range musical structure while preserving fine-grained acoustic details and enabling efficient autoregressive modeling; (4) HeartMuLa, an LLM-based song generation model capable of synthesizing high-fidelity music under rich, user-controllable conditions (e.g., textual style descriptions, lyrics, and reference audio). In addition, it provides two specialized modes: (i) fine-grained musical attribute control, which allows users to specify the style of different song sections (e.g., intro, verse, chorus) using natural language prompts; and (ii) short, engaging music generation, which is suitable as background music for short videos. Lastly, HeartMuLa improves significantly when scaled to 7B parameters. For the first time, we show that a Suno-level, commercial-grade system can be reproduced using academic-scale data and GPU resources. We expect these foundation models to serve as strong baselines for future research and to facilitate practical applications in multimodal content production.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。