用音乐大模型提升音频指纹抗干扰能力,适配短视频变声压损场景。
Robust Neural Audio Fingerprinting using Music Foundation Models
- 以音乐基础模型为骨干,替代传统训练方式
- 在时间拉伸、压缩等12种干扰下仍保持高匹配率
- 适合音乐版权管理与海量曲库检索场景
现代媒体平台(如TikTok)上充斥着失真、压缩和篡改的音乐内容,亟需更鲁棒的音频指纹技术来识别音乐来源。本文提出新型神经音频指纹方法,实现两大改进:一是采用预训练音乐基础模型(如MuQ、MERT)作为神经网络主干;二是扩展数据增强策略,在时间拉伸、音高调制、压缩和滤波等多种音频扰动下训练模型。系统评估表明,基于音乐基础模型提取的指纹显著优于从零开始训练或在非音乐音频上预训练的模型。段级评估进一步验证其精准定位匹配片段的能力,对曲库管理具有重要实用价值。
原文摘要 · Abstract (English)
The proliferation of distorted, compressed, and manipulated music on modern media platforms like TikTok motivates the development of more robust audio fingerprinting techniques to identify the sources of musical recordings. In this paper, we develop and evaluate new neural audio fingerprinting techniques with the aim of improving their robustness. We make two contributions to neural fingerprinting methodology: (1) we use a pretrained music foundation model as the backbone of the neural architecture and (2) we expand the use of data augmentation to train fingerprinting models under a wide variety of audio manipulations, including time streching, pitch modulation, compression, and filtering. We systematically evaluate our methods in comparison to two state-of-the-art neural fingerprinting models: NAFP and GraFPrint. Results show that fingerprints extracted with music foundation models (e.g., MuQ, MERT) consistently outperform models trained from scratch or pretrained on non-musical audio. Segment-level evaluation further reveals their capability to accurately localize fingerprint matches, an important practical feature for catalog management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。