arXiv:2509.14052cs.SDeess.SP2025-09被引 5

让伴奏模型在真实歌声上也能用,突破了传统方法的泛化瓶颈。

AnyAccomp: Generalizable Accompaniment Generation via Quantized Melodic Bottleneck

  • 用离散旋律编码分离音色干扰,提取鲁棒旋律特征
  • 在真实人声和独奏乐器上表现远超现有模型
  • 适合音乐创作、跨乐器伴奏生成等实际应用

演唱伴奏生成(SAG)是为给定纯净人声输入生成伴奏音乐的过程。然而,现有方法依赖于源分离后的人声输入,容易过拟合于分离产生的伪影,导致训练与测试条件严重不一致,无法处理真实世界中的纯净人声。本文提出AnyAccomp框架,通过将伴奏生成与源相关伪影解耦来解决此问题。首先利用音高图与VQ-VAE构建量化旋律瓶颈,提取离散且与音色无关的核心旋律表示;随后由流匹配模型基于这些鲁棒编码生成伴奏。实验表明,AnyAccomp在分离人声基准上表现相当,但在真实录音室人声和独奏乐器的泛化测试集上显著优于基线模型。这标志着泛化能力的质变,首次实现对乐器伴奏的有效生成,而此前模型完全失效,为更通用的音乐协同创作工具铺平道路。演示音频与代码:https://anyaccomp.github.io

原文摘要 · Abstract (English)

Singing Accompaniment Generation (SAG) is the process of generating instrumental music for a given clean vocal input. However, existing SAG techniques use source-separated vocals as input and overfit to separation artifacts. This creates a critical train-test mismatch, leading to failure on clean, real-world vocal inputs. We introduce AnyAccomp, a framework that resolves this by decoupling accompaniment generation from source-dependent artifacts. AnyAccomp first employs a quantized melodic bottleneck, using a chromagram and a VQ-VAE to extract a discrete and timbre-invariant representation of the core melody. A subsequent flow-matching model then generates the accompaniment conditioned on these robust codes. Experiments show AnyAccomp achieves competitive performance on separated-vocal benchmarks while significantly outperforming baselines on generalization test sets of clean studio vocals and, notably, solo instrumental tracks. This demonstrates a qualitative leap in generalization, enabling robust accompaniment for instruments - a task where existing models completely fail - and paving the way for more versatile music co-creation tools. Demo audio and code: https://anyaccomp.github.io

伴奏生成泛化能力VQ-VAE音乐AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。