arXiv:2508.11609cs.SDcs.AI2025-08被引 1

用3秒音频生成唯一指纹,抗各种失真且效果领先

Pretrained Conformers for Audio Fingerprinting and Retrieval

  • 基于自监督对比学习训练音视频编码器
  • 仅需3秒音频即生成可检索嵌入,性能超越现有方法
  • 对时间错位和噪声等失真几乎免疫,适合实际应用

由于能够同时捕捉局部与全局交互,Conformer 在语音处理中表现优异。本文采用自监督对比学习框架,训练基于 Conformer 的编码器,使其能为短时音频片段生成唯一嵌入,并在未见数据上具有良好泛化能力。实验表明,仅需3秒音频即可生成嵌入,在音频检索任务中达到当前最优性能。模型几乎完全不受时间错位影响,在噪声、混响或极端时间拉伸等音频失真场景下也表现卓越。代码与模型已公开,使用常见且免费的数据集进行训练与测试,结果易于复现。

原文摘要 · Abstract (English)

Conformers have shown great results in speech processing due to their ability to capture both local and global interactions. In this work, we utilize a self-supervised contrastive learning framework to train conformer-based encoders that are capable of generating unique embeddings for small segments of audio, generalizing well to previously unseen data. We achieve state-of-the-art results for audio retrieval tasks while using only 3 seconds of audio to generate embeddings. Our models are almost completely immune to temporal misalignments and achieve state-of-the-art results in cases of other audio distortions such as noise, reverb or extreme temporal stretching. Code and models are made publicly available and the results are easy to reproduce as we train and test using popular and freely available datasets of different sizes.

音频指纹自监督学习音频检索Conformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。