arXiv:2602.10058cs.SDcs.LG2026-02中稿 · ICASSP 2026被引 2

评估音乐生成中解耦表征的实际效果,发现其语义与预期不符。

Evaluating Disentangled Representations for Controllable Music Generation

  • 用探测框架分析多种无监督解耦策略的效果。
  • 发现嵌入语义与设计意图不一致,解耦不彻底。
  • 适合关注可控音乐生成可靠性的研究者参考。

近期音乐生成方法依赖解耦表征(如结构与音色、局部与全局),以实现可控合成。然而这些嵌入的底层特性仍缺乏深入探索。本文通过超越常规下游任务的探测框架,评估一组音乐音频模型中的解耦表征在可控生成中的表现。所选模型涵盖归纳偏置、数据增强、对抗目标及分阶段训练等多种无监督解耦策略,并进一步分离具体策略以分析其影响。分析覆盖信息量、等变性、不变性和解耦性四个关键维度,跨数据集、任务和受控变换进行评估。结果揭示嵌入的预期语义与实际语义存在不一致,表明当前策略难以实现真正解耦,促使我们重新审视音乐生成中的可控性设计方式。

原文摘要 · Abstract (English)

Recent approaches in music generation rely on disentangled representations, often labeled as structure and timbre or local and global, to enable controllable synthesis. Yet the underlying properties of these embeddings remain underexplored. In this work, we evaluate such disentangled representations in a set of music audio models for controllable generation using a probing-based framework that goes beyond standard downstream tasks. The selected models reflect diverse unsupervised disentanglement strategies, including inductive biases, data augmentations, adversarial objectives, and staged training procedures. We further isolate specific strategies to analyze their effect. Our analysis spans four key axes: informativeness, equivariance, invariance, and disentanglement, which are assessed across datasets, tasks, and controlled transformations. Our findings reveal inconsistencies between intended and actual semantics of the embeddings, suggesting that current strategies fall short of producing truly disentangled representations, and prompting a re-examination of how controllability is approached in music generation.

音乐生成解耦表征可控合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。