用遗传编程自动组合音频特征,提升音乐标签准确率并保持可解释性。
Constructing Composite Features for Interpretable Music-Tagging
- 通过遗传编程数学组合基础特征,生成可解释的复合特征。
- 在MTG-Jamendo和GTZAN数据集上优于现有系统,前几百次迭代即见效。
- 结果揭示有效特征交互模式,适合需要可解释性的音乐分析场景。
融合多个音频特征可提升音乐标签性能,但主流深度学习融合方法缺乏可解释性。为此,我们提出一种遗传编程(GP)流程,通过数学方式自动演化复合特征,捕捉特征间协同作用的同时保持可解释性。该方法在代表性和性能上媲美深度特征融合,却无需牺牲可解释性。在MTG-Jamendo与GTZAN数据集上的实验表明,无论基础特征抽象层级如何,本方法均持续优于当前最优系统。值得注意的是,多数性能提升出现在前几百次GP评估内,说明在较小搜索预算下即可发现有效特征组合。最优演化表达式包含线性、非线性及条件形式,且高性能解普遍具有低复杂度,符合简约性原则。进一步分析显示,某些特定交互与变换对标签任务更有效,这些洞察在黑箱深度模型中难以获得。
原文摘要 · Abstract (English)
Combining multiple audio features can improve the performance of music tagging, but common deep learning-based feature fusion methods often lack interpretability. To address this problem, we propose a Genetic Programming (GP) pipeline that automatically evolves composite features by mathematically combining base music features, thereby capturing synergistic interactions while preserving interpretability. This approach provides representational benefits similar to deep feature fusion without sacrificing interpretability. Experiments on the MTG-Jamendo and GTZAN datasets demonstrate consistent improvements compared to state-of-the-art systems across base feature sets at different abstraction levels. It should be noted that most of the performance gains are noticed within the first few hundred GP evaluations, indicating that effective feature combinations can be identified under modest search budgets. The top evolved expressions include linear, nonlinear, and conditional forms, with various low-complexity solutions at top performance aligned with parsimony pressure to prefer simpler expressions. Analyzing these composite features further reveals which interactions and transformations tend to be beneficial for tagging, offering insights that remain opaque in black-box deep models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。