评估主流音视频嵌入模型对音色感知的捕捉能力,发现LAION-CLAP表现最优。
Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
- 构建双实验验证语言-音频嵌入对音色感知的对齐效果
- LAION-CLAP在乐器与效果音中对音色感知对齐最一致
- 混响引起的音色变化比均衡更易被模型捕捉
理解语言与声音之间的关系对音乐信息检索、文本引导音乐生成和音频描述等应用至关重要。核心在于联合语言-音频嵌入空间,将文本描述与听觉内容映射到共享表征。尽管多模态嵌入模型如MS-CLAP、LAION-CLAP、MuQ-MuLan和OpenFLAM在语言-音频对齐上表现优异,但其与人类对音色(包括明亮度、粗糙度、温暖感等多维属性)感知的对应关系仍缺乏深入探索。本文通过两项互补实验评估这些模型在捕捉感知音色语义方面的能力。结果表明,LAION-CLAP在乐器声音和描述条件下的音频效果中表现出相对更强且一致的对齐能力。然而整体对齐程度有限,说明当前联合嵌入仅部分捕捉了音色感知。此外,总体而言,混响引起的音色语义比均衡引起的更一致地被编码。
原文摘要 · Abstract (English)
Understanding and modeling the relationship between language and sound are essential for applications such as music information retrieval, text-guided music generation, and audio captioning. Central to these tasks are joint language-audio embedding spaces, which map textual descriptions and auditory content into a shared representation. Although multimodal embedding models such as MS-CLAP, LAION-CLAP, MuQ-MuLan, and OpenFLAM have shown strong performance in language-audio alignment, their correspondence to human perception of timbre, a multifaceted attribute encompassing qualities such as brightness, roughness, and warmth, remains under-explored. In this paper, we evaluate these joint language-audio embedding models in terms of their ability to capture perceptual timbre semantics. Across two complementary experiments, we find that LAION-CLAP shows relatively strong and consistent alignment with human-perceived timbre semantics across both instrumental sounds and descriptor-conditioned audio effects. At the same time, the overall strength of this alignment remains limited, suggesting that current joint language-audio embeddings capture perceptual timbre semantics only partially. We also observe that, overall, reverb-induced timbre semantics are more consistently encoded than equalization-induced timbre semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。