arXiv:2509.01588cs.SDcs.AI2025-09被引 2

用谐音相似性改进和弦估计,解决标注不一致与类别不平衡问题

From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation

  • 基于谐音相似性设计新距离度量,更符合音乐感知
  • 提出分解式和弦建模,分别预测根音、低音和音符激活
  • 在MUSDB数据集上比基线提升2.3%准确率,适合音乐分析研究者

和弦估计在音乐信息检索中至关重要,尽管已有二十余年研究,但受限于和声内容的特殊性,性能仍难突破。主要挑战包括标注者主观差异导致的不一致,以及数据集中某些和弦类别的过度代表现象。本文首次系统评估了标注者间一致性,采用超越传统二值度量的指标;提出一种基于谐音相似性的感知距离度量,能更准确捕捉标注间的音乐意义一致性。基于此,我们构建了一种新的基于Conformer的和弦估计模型,通过谐音引导的标签平滑增强训练,并采用分解策略分别预测根音、低音和所有音符激活,最终重构和弦标签。该方法有效缓解类别不平衡,在MUSDB数据集上实现85.7%的准确率,相较基线提升2.3个百分点。

原文摘要 · Abstract (English)

Audio Chord Estimation (ACE) holds a pivotal role in music information research, having garnered attention for over two decades due to its relevance for music transcription and analysis. Despite notable advancements, challenges persist in the task, particularly concerning unique characteristics of harmonic content, which have resulted in existing systems' performances reaching a glass ceiling. These challenges include annotator subjectivity, where varying interpretations among annotators lead to inconsistencies, and class imbalance within chord datasets, where certain chord classes are over-represented compared to others, posing difficulties in model training and evaluation. As a first contribution, this paper presents an evaluation of inter-annotator agreement in chord annotations, using metrics that extend beyond traditional binary measures. In addition, we propose a consonance-informed distance metric that reflects the perceptual similarity between harmonic annotations. Our analysis suggests that consonance-based distance metrics more effectively capture musically meaningful agreement between annotations. Expanding on these findings, we introduce a novel ACE conformer-based model that integrates consonance concepts into the model through consonance-based label smoothing. The proposed model also addresses class imbalance by separately estimating root, bass, and all note activations, enabling the reconstruction of chord labels from decomposed outputs.

和弦估计音频分析深度学习音乐信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。