构建首个电吉他音色语义标注数据集,支持音色可控生成。
A Semantic Timbre Dataset for the Electric Guitar
- 基于19个语义描述符系统标注电吉他音色,关联感知维度与标签
- 训练的变分自编码器可实现语义描述符间的平滑音色插值
- 适合音色控制、生成式音频和人机感知研究者使用
理解与操控音色是音频合成的核心,但因缺乏将听觉感知音色维度与语义描述关联的标注数据集,该领域在机器学习中仍进展有限。本文提出语义音色数据集(Semantic Timbre Dataset),包含经系统标注的单音电吉他音色样本,每个样本配有19个语义音色描述符及其对应强度。这些描述符源自对物理与虚拟吉他效果单元的定性分析,并应用于纯净吉他音色。该数据集连接了音色感知与机器学习表征,支持音色控制与语义音频生成的学习。通过在潜在空间训练变分自编码器(VAE),并结合人类听觉判断与描述符分类器评估,结果表明该模型有效捕捉音色结构,并实现描述符间的平滑过渡。数据集、代码与评估协议已公开,以推动音色感知生成式AI研究。
原文摘要 · Abstract (English)
Understanding and manipulating timbre is central to audio synthesis, yet this remains under-explored in machine learning due to a lack of annotated datasets linking perceptual timbre dimensions to semantic descriptors. We present the Semantic Timbre Dataset, a curated collection of monophonic electric guitar sounds, each labeled with one of 19 semantic timbre descriptors and corresponding magnitudes. These descriptors were derived from a qualitative analysis of physical and virtual guitar effect units and applied systematically to clean guitar tones. The dataset bridges perceptual timbre and machine learning representations, supporting learning for timbre control and semantic audio generation. We validate the dataset by training a variational autoencoder (VAE) on its latent space and evaluating it using human perceptual judgments and descriptor classifiers. Results show that the VAE captures timbral structure and enables smooth interpolation across descriptors. We release the dataset, code, and evaluation protocols to support timbre-aware generative AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。