arXiv:2606.24307cs.SDcs.AI2026-06被引 1

让音乐生成像乐器一样实时互动,无需等待。

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

论文配图:Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation
图 1 · 摘自论文原文
  • 用流式自回归潜空间蒸馏,边生成边优化。
  • 单步加速生成,实时因子极低,音质保真度高。
  • 适合现场音乐人与AI即时共创,无需训练数据。

实时互动音乐与现场表演依赖于即时的人类表达,但现有生成式音乐AI因推理延迟过高和离线生成范式而难以融入该领域。为使先锋音乐人获得新型交互创作媒介,需将静态模型转化为动态可演奏的工具。本文提出一种框架,通过在流式自回归潜空间中进行蒸馏,实现低延迟下的结构连贯性。该方法仅使用提示输入,无需昂贵的音频-潜空间配对数据集,即可在线合成教师引导的分块轨迹。为保障乐器级音质,引入音乐感知的一致性目标,结合潜空间、频谱与时间差损失,保留音色、瞬态与节奏稳定性。通过参数高效适配,蒸馏显著减少生成步骤,达成低实时因子。关键在于系统以连续自回归流运行,可无缝接收动态人类输入,用户能即时引导乐曲进程而不中断音频流。本工作将文本到音乐生成模型从被动的‘提示-等待’系统重构为响应式乐器,开启人机共创音乐的新可能。

原文摘要 · Abstract (English)

Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from this domain due to its prohibitive inference latency and offline rendering paradigm. To provide pioneer musicians with a novel medium for interactive composition, we should fundamentally change these static models into dynamic, playable instruments. In this paper, we propose a framework that bridges this gap. To achieve the low latency required for live interaction without sacrificing structural coherence, we formulate distillation within a streaming autoregressive latent space. Our approach gets rid of the need for expensive paired audio-latent datasets by utilizing prompt-only inputs to synthesize teacher-guided, chunk-wise trajectories on the fly. Because live instruments require high acoustic fidelity, we introduce music-aware consistency objectives, which combine latent, spectral, and temporal-difference losses, to preserve crucial qualities like timbre, transients, and rhythmic stability during accelerated single-step streaming generation. Implemented via parameter-efficient adaptation, our distillation reduces generation steps to achieve a low real-time factor. Crucially, by operating as a continuous autoregressive stream, the system can seamlessly assimilate dynamic human inputs on the fly, allowing users to instantly steer the musical trajectory without interrupting the audio flow. Ultimately, this work recontextualizes generative text-to-music models not as passive prompt-and-wait systems, but as responsive instruments, opening new frontiers for live human-AI musical co-creation.

音乐生成实时交互蒸馏自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。