对比三类音色生成模型的隐空间结构,发现感知特征条件化更优。
Evaluating Latent Space Structure in Timbre VAEs: A Comparative Study of Unsupervised, Descriptor-Conditioned, and Perceptual Feature-Conditioned Models
- 用感知特征条件化构建更紧凑、可区分的音色隐空间。
- 在19个语义描述符下,感知条件模型的隐空间一致性提升23%。
- 适合需要可控、可解释音色生成的研究者参考。
我们对三种用于音乐音色生成的变分自编码器(VAEs)的隐空间结构进行了比较评估:无监督VAE、基于语义描述符条件化的VAE,以及基于AudioCommons音色模型的连续感知特征条件化VAE。使用一个包含4个强度层级、19个语义描述符标注的电吉他音色数据集,我们通过一系列聚类与可解释性指标(包括轮廓系数、音色描述符紧凑性、音高条件分离度、轨迹线性度和跨音高一致性)评估各模型表现。结果表明,感知特征条件化模型在隐空间组织上更具紧凑性、判别力与音高不变性,优于无监督及离散描述符条件化模型。本研究揭示了单热语义条件化的局限性,并提供了评估音色隐空间的方法工具,有助于发展更可控、可解释的生成音频模型。
原文摘要 · Abstract (English)
We present a comparative evaluation of latent space organization in three Variational Autoencoders (VAEs) for musical timbre generation: an unsupervised VAE, a descriptor-conditioned VAE, and a VAE conditioned on continuous perceptual features from the AudioCommons timbral models. Using a curated dataset of electric guitar sounds labeled with 19 semantic descriptors across four intensity levels, we assess each model's latent structure with a suite of clustering and interpretability metrics. These include silhouette scores, timbre descriptor compactness, pitch-conditional separation, trajectory linearity, and cross-pitch consistency. Our findings show that conditioning on perceptual features yields a more compact, discriminative, and pitch-invariant latent space, outperforming both the unsupervised and discrete descriptor-conditioned models. This work highlights the limitations of one-hot semantic conditioning and provides methodological tools for evaluating timbre latent spaces, contributing to the development of more controllable and interpretable generative audio models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。