给实时伴奏模型装上节拍器,解决节奏漂移问题。
Silent Metronome: Rhythmic Grounding for Live Music Accompaniment

- 用周期性信号编码节拍与小节相位,作为独立条件输入。
- 节拍对齐性能提升3.2倍,超越带一秒前瞻的非因果模型。
- 适合需要精准节奏同步的实时音乐生成任务。
实时伴奏模型在严格因果条件下生成音乐,需根据自身不完美的历史推断节拍、小节和节拍相位,导致误差快速累积并产生可听的节奏漂移。模型虽有‘耳朵’,却无时间参考,当听到模糊或不准确的音乐时,输出将出错。本文提出 Silent Metronome(SiMe),通过将节拍内相位与小节内相位编码为周期函数,结合速度与拍号,以独立条件通道提供时间参考。该参考与生成音频无关,不会漂移。额外设计的辅助头(包括预测自身未来标记的新头)优化潜在表示。使用真实标注的节拍信息,节拍对齐性能相比严格因果基线提升3.2倍,超过获得一秒钟前瞻的非因果参考模型。输入与伴奏间的连贯性保持在参考点的一点之内。结果表明,流式伴奏系统应像人类乐团一样共享节奏信号,而非自行推断。
原文摘要 · Abstract (English)
Live accompaniment models generate music for an incoming audio stream, committing to each output frame before hearing what comes next. In this strictly causal setting the model must infer tempo, meter, and metrical phase from its own imperfect past, whereby compounding errors quickly become audible as rhythmic drift. Put simply, the model has ears but no temporal reference, so when the ears hear imperfect, ambiguous music, the model will produce a flawed output. We propose Silent Metronome (SiMe), which gives it the temporal reference, encoding the phase within the beat and within the bar as periodic functions, pairing them with tempo and time signature, and supplying the result as a separate conditioning channel. Because this reference is independent of the generated audio, it cannot drift. Complementary auxiliary heads shape the latent representation, including a novel head that predicts the model's own future tokens. With the metrical signal taken from ground-truth annotations, beat alignment improves by a factor of 3.2 over the strictly causal baseline and surpasses a non-causal reference granted a full second of look-ahead. Coherence between input and accompaniment stays within a single point of that reference. These results suggest that streaming accompaniment systems should treat rhythm as a signal to be shared, as human ensembles do, rather than inferred.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。