arXiv:2606.30369cs.SD2026-06

用深度模型预测音色特征,让声音合成器评估更可解释。

Predicting Timbre Traits for Interpretable Assessment of Musical Sound Synthesizers

论文配图:Predicting Timbre Traits for Interpretable Assessment of Musical Sound Synthesizers
图 1 · 摘自论文原文
  • 用CLAP+浅层网络预测20种音色描述,基于人类评分训练。
  • 预测结果与人类平均评分相关性达r=0.66(p<0.001)。
  • 能定位合成音色问题,适合音色设计与优化研究者。

衡量神经音频合成器性能通常采用基于分布的指标,如弗雷歇音频距离(FAD)。尽管该指标与人类感知相关,但仅能排序不同方法,缺乏可解释性。本文提出一种深度音色特征预测模型,由预训练音频嵌入(CLAP)和浅层可学习组件构成。该组件在RWC乐器数据库上训练,使用人类对31种乐器20种音色描述(如木质、打击感、轰鸣等)的评分。模型与平均人类评分的相关性达r=0.66(p<0.001)。进一步应用于TokenSynth合成器评估:首先,不同条件下的生成音色在预测值上的平均绝对误差(MAE)与基于RWC参考的FAD排名一致,表明其具备分布级信息等效性;其次,因模型可定性分析单个声音,能识别需改进的生成样本及具体音色维度,实现精准优化。

原文摘要 · Abstract (English)

Measuring neural audio synthesizers' performance is now routinely conducted using distribution based metrics such as the Fréchet Audio Distance (FAD). Although this metric can be correlated with human perception, it offers limited interpretability beyond ranking different approaches. In this paper, we introduce a deep neural timbre trait predictor composed of a pretrained audio neural embedding (CLAP), and a shallow learnable component. The latter is trained using the RWC musical instrument database and human judgments of 20 timbre descriptions (e.g., woody, percussive, rumbling, etc.) for 31 instruments. The resulting model shows strong correlation with average human ratings (r = 0.66, p < 0.001). We then demonstrate the benefit of this predictor for evaluating the performance of TokenSynth, a neural sound synthesizer. First, the Mean Absolute Error (MAE) computed over the set of generated sounds under different conditioning modalities of the model provides the same ranking as a FAD computed with the RWC database as a reference, suggesting that the proposed predictors are able to provide equivalent information on a distributional basis. Second, because the model is able to qualitatively analyze isolated sounds, we can determine which generated sounds could be improved and identify specific timbral dimensions that need adjustment.

音色预测可解释性声音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。