提升音频指纹对真实降噪的鲁棒性,效果更接近实际使用场景。
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
- 利用音乐信号特性和真实房间声学构建自监督训练策略
- 发现三元组损失在音频指纹任务中优于主流的NT-Xent损失
- 多正样本训练对不同损失函数影响显著不同,需针对性设计
音频指纹(AFP)通过提取紧凑表示实现未知音频内容识别,这些表示应能抵抗常见音频退化。神经网络方法通常采用度量学习,其表示质量受监督信号和损失函数影响。然而,现有研究在训练时不切实际地模拟真实音频退化,导致监督信号不足。尽管已有多种现代度量学习方法,当前神经AFP仍主要依赖NT-Xent损失,未充分探索近期进展或经典替代方案。本文提出一系列最佳实践,通过利用音乐信号特性与真实房间声学增强自监督能力。首次系统评估多种度量学习方法在AFP中的表现,发现三元组损失的自监督变体性能最优。结果还表明,每个锚点使用多个正样本在不同损失函数下影响迥异。基于此,我们的方法在大规模合成退化数据集及真实世界麦克风采集数据集上均达到当前最优性能。
原文摘要 · Abstract (English)
Audio fingerprinting (AFP) allows the identification of unknown audio content by extracting compact representations, termed audio fingerprints, that are designed to remain robust against common audio degradations. Neural AFP methods often employ metric learning, where representation quality is influenced by the nature of the supervision and the utilized loss function. However, recent work unrealistically simulates real-life audio degradation during training, resulting in sub-optimal supervision. Additionally, although several modern metric learning approaches have been proposed, current neural AFP methods continue to rely on the NT-Xent loss without exploring the recent advances or classical alternatives. In this work, we propose a series of best practices to enhance the self-supervision by leveraging musical signal properties and realistic room acoustics. We then present the first systematic evaluation of various metric learning approaches in the context of AFP, demonstrating that a self-supervised adaptation of the triplet loss yields superior performance. Our results also reveal that training with multiple positive samples per anchor has critically different effects across loss functions. Our approach is built upon these insights and achieves state-of-the-art performance on both a large, synthetically degraded dataset and a real-world dataset recorded using microphones in diverse music venues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。