用注意力机制捕捉音乐时序关键片段,提升流派分类准确率。
Attention-guided Spectrogram Sequence Modeling with CNNs for Music Genre Classification
- 通过CNN与多头注意力处理频谱序列,聚焦重要时间片段。
- 有效识别流派特异性特征,提升分类性能。
- 适合音乐推荐与人感匹配的研究者参考。
音乐流派分类是音乐推荐系统、生成算法和文化分析的核心组成部分。本文提出一种基于注意力机制的时序特征建模方法,通过卷积神经网络(CNN)和多头注意力层处理频谱序列,捕捉每首音乐中最关键的时序片段,构建用于流派识别的独特“签名”。该方法不仅提升了分类准确率,还揭示了流派特有特征,可直观映射到听觉感知。研究结果为个性化音乐推荐系统提供新思路,揭示跨流派的共性与差异,与人类音乐直觉高度契合,弥合了技术分类与人类体验之间的鸿沟。
原文摘要 · Abstract (English)
Music genre classification is a critical component of music recommendation systems, generation algorithms, and cultural analytics. In this work, we present an innovative model for classifying music genres using attention-based temporal signature modeling. By processing spectrogram sequences through Convolutional Neural Networks (CNNs) and multi-head attention layers, our approach captures the most temporally significant moments within each piece, crafting a unique "signature" for genre identification. This temporal focus not only enhances classification accuracy but also reveals insights into genre-specific characteristics that can be intuitively mapped to listener perceptions. Our findings offer potential applications in personalized music recommendation systems by highlighting cross-genre similarities and distinctiveness, aligning closely with human musical intuition. This work bridges the gap between technical classification tasks and the nuanced, human experience of genre.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。