arXiv:2508.04651cs.SDcs.HC2025-08被引 21

实时生成可交互的音乐,让用户用文字或音频即时控制风格。

Live Music Models

  • 通过文本或音频提示实时控制音乐风格,实现连续流式生成。
  • 参数更少但音乐质量超越现有开源模型,支持实时演奏。
  • 适合音乐人、创作者用于现场演出中的智能辅助创作。

我们提出一类新型生成式音乐模型——实时音乐模型,可在实时流中生成连续音乐,并支持用户同步控制。我们发布了 Magenta RealTime,一个开源权重的实时音乐模型,可通过文本或音频提示控制音色风格。在音乐质量自动评估指标上,尽管参数量更少,Magenta RealTime 仍优于其他开源音乐生成模型,并首次实现真正的实时生成能力。我们还发布了 Lyria RealTime,一个基于 API 的模型,提供更丰富的控制选项和对最强大模型的访问,覆盖广泛提示类型。这些模型展现了以人类为中心的 AI 音乐创作新范式,强调在实时音乐表演中的人机协同。

原文摘要 · Abstract (English)

We introduce a new class of generative models for music called live music models that produce a continuous stream of music in real-time with synchronized user control. We release Magenta RealTime, an open-weights live music model that can be steered using text or audio prompts to control acoustic style. On automatic metrics of music quality, Magenta RealTime outperforms other open-weights music generation models, despite using fewer parameters and offering first-of-its-kind live generation capabilities. We also release Lyria RealTime, an API-based model with extended controls, offering access to our most powerful model with wide prompt coverage. These models demonstrate a new paradigm for AI-assisted music creation that emphasizes human-in-the-loop interaction for live music performance.

实时生成人机交互音乐生成AI作曲

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。