让音乐自动打标签更透明,用多模态特征解释预测结果
Semantic-Aware Interpretable Multimodal Music Auto-Tagging
- 分组融合信号处理、深度学习等多源特征
- 通过聚类和算法分配每组特征权重
- 适合需要理解标签依据的研究者与开发者
音乐自动打标签对大规模数字音乐库的组织与发现至关重要。尽管基础模型在此任务上表现优异,但其输出往往缺乏可解释性,限制了研究人员与用户对其的信任与使用。本文提出一种可解释的音乐自动打标签框架,利用来自信号处理、深度学习、本体工程和自然语言处理的、具有音乐意义的多模态特征组。为增强可解释性,我们基于语义对特征进行聚类,并采用期望最大化算法,根据各特征组在打标过程中的贡献赋予不同权重。该方法在保持竞争性性能的同时,提供了对决策过程的深层理解,为更透明、以用户为中心的音乐打标系统铺平道路。
原文摘要 · Abstract (English)
Music auto-tagging is essential for organizing and discovering music in extensive digital libraries. While foundation models achieve exceptional performance in this domain, their outputs often lack interpretability, limiting trust and usability for researchers and end-users alike. In this work, we present an interpretable framework for music auto-tagging that leverages groups of musically meaningful multimodal features, derived from signal processing, deep learning, ontology engineering, and natural language processing. To enhance interpretability, we cluster features semantically and employ an expectation maximization algorithm, assigning distinct weights to each group based on its contribution to the tagging process. Our method achieves competitive tagging performance while offering a deeper understanding of the decision-making process, paving the way for more transparent and user-centric music tagging systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。