arXiv:2507.07764cs.SDeess.AS2025-07中稿 · ISMIR 2025被引 7

用人类听觉数据测试音频表示与音色相似性的对齐度,发现图像风格迁移思路的嵌入效果最好。

Assessing the Alignment of Audio Representations with Timbre Similarity Ratings

  • 用距离值和排序对比方法评估多种音频表示与人类音色感知的匹配度
  • 来自CLAP和新声学匹配模型的风格嵌入表现最佳,超越其他15种表示
  • 适合研究音频感知、音色建模或跨模态表示学习的研究者参考

心理声学中的“音色空间”通过多维缩放将乐器声音的感知相似性映射到低维嵌入,但存在可扩展性差和泛化能力弱的问题。近期在音频质量评估及图像相似性任务中,深度学习已证明能生成与人类感知高度一致的嵌入,且不受上述限制。尽管现有音色相似性人工评分数据量有限(334个音频样本,2,614组配对评分),仍可用于测试音频模型。本文引入新指标,通过比较嵌入距离的绝对值与排序,评估多种音频表示与人类音色相似性判断的对齐程度。实验涵盖三种信号处理表示、十二种预训练模型提取的表示,以及三种由新型声学匹配模型生成的表示。结果显示,受图像风格迁移启发的风格嵌入(来自CLAP模型和新模型)显著优于其他表示,展现出建模音色相似性的潜力。

原文摘要 · Abstract (English)

Psychoacoustical so-called "timbre spaces" map perceptual similarity ratings of instrument sounds onto low-dimensional embeddings via multidimensional scaling, but suffer from scalability issues and are incapable of generalization. Recent results from audio (music and speech) quality assessment as well as image similarity have shown that deep learning is able to produce embeddings that align well with human perception while being largely free from these constraints. Although the existing human-rated timbre similarity data is not large enough to train deep neural networks (2,614 pairwise ratings on 334 audio samples), it can serve as test-only data for audio models. In this paper, we introduce metrics to assess the alignment of diverse audio representations with human judgments of timbre similarity by comparing both the absolute values and the rankings of embedding distances to human similarity ratings. Our evaluation involves three signal-processing-based representations, twelve representations extracted from pre-trained models, and three representations extracted from a novel sound matching model. Among them, the style embeddings inspired by image style transfer, extracted from the CLAP model and the sound matching model, remarkably outperform the others, showing their potential in modeling timbre similarity.

音色建模深度表示感知对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。