arXiv:2511.07118cs.LGcs.AI2025-11

通过优化属性变换提升音乐生成的可控性与潜在空间正则化

On the Joint Minimization of Regularization Loss Functions in Deep Variational Bayesian Methods for Attribute-Controlled Symbolic Music Generation

  • 引入属性变换缓解变分瓶颈中正则化项的冲突
  • 实验证明可同时实现属性可控与潜变量分布约束
  • 适合关注音乐生成控制力与模型稳定性的研究者

显式潜在变量模型为数据合成提供了灵活而强大的框架,支持对生成因素的精确控制。通过从可计算的概率密度函数中采样潜在变量并加以约束,可在潜在空间中实现连续且语义丰富的输出探索。结构化潜在表示通常通过联合最小化正则化损失函数获得。在变分信息瓶颈模型中,重建损失与KL散度(KLD)常与辅助属性正则化(AR)损失线性组合。然而,平衡KLD与AR极为困难:当KLD主导时,生成模型缺乏可控性;当AR主导时,随机编码器会违背标准正态先验。本文在符号音乐生成中显式控制连续音乐属性的背景下探索该权衡问题,发现现有方法难以同时最小化两类正则化目标,而合适的属性变换可有效实现可控性与目标潜变量维度的正则化。

原文摘要 · Abstract (English)

Explicit latent variable models provide a flexible yet powerful framework for data synthesis, enabling controlled manipulation of generative factors. With latent variables drawn from a tractable probability density function that can be further constrained, these models enable continuous and semantically rich exploration of the output space by navigating their latent spaces. Structured latent representations are typically obtained through the joint minimization of regularization loss functions. In variational information bottleneck models, reconstruction loss and Kullback-Leibler Divergence (KLD) are often linearly combined with an auxiliary Attribute-Regularization (AR) loss. However, balancing KLD and AR turns out to be a very delicate matter. When KLD dominates over AR, generative models tend to lack controllability; when AR dominates over KLD, the stochastic encoder is encouraged to violate the standard normal prior. We explore this trade-off in the context of symbolic music generation with explicit control over continuous musical attributes. We show that existing approaches struggle to jointly minimize both regularization objectives, whereas suitable attribute transformations can help achieve both controllability and regularization of the target latent dimensions.

音乐生成变分推断潜在空间控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。