arXiv:2509.18620cs.SDcs.IR2025-09

用生成模型模拟音频指纹,实现大规模无音频评估。

Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation

  • 训练可逆流模型生成逼近真实指纹分布的合成指纹。
  • 合成干扰项使检索性能评估与真实数据趋势一致。
  • 无需真实音频即可评估超大规模系统扩展性,适合算法评测者。

真实规模的音频指纹评估受限于大型公开音乐数据库的稀缺。本文提出一种无需音频的方法,通过预训练神经音频指纹系统提取的嵌入,训练修正流模型以合成逼近真实指纹分布的潜在指纹。生成的合成指纹作为逼真干扰项,可在不依赖额外音频的情况下实现大规模检索性能模拟。通过对比分布验证合成指纹的真实性,并在多个前沿指纹框架上,以合成干扰项扩充真实参考库进行基准测试,结果显示合成干扰项所得的缩放趋势与真实干扰项高度吻合。最终,将合成干扰项数据库扩展至极大规模,提供不依赖音频语料的系统可扩展性实用度量。

原文摘要 · Abstract (English)

The evaluation of audio fingerprinting at a realistic scale is limited by the scarcity of large public music databases. We present an audio-free approach that synthesises latent fingerprints which approximate the distribution of real fingerprints. Our method trains a Rectified Flow model on embeddings extracted by pre-trained neural audio fingerprinting systems. The synthetic fingerprints generated using our system act as realistic distractors and enable the simulation of retrieval performance at a large scale without requiring additional audio. We assess the fidelity of synthetic fingerprints by comparing the distributions to real data. We further benchmark the retrieval performances across multiple state-of-the-art audio fingerprinting frameworks by augmenting real reference databases with synthetic distractors, and show that the scaling trends obtained with synthetic distractors closely track those obtained with real distractors. Finally, we scale the synthetic distractor database to model retrieval performance for very large databases, providing a practical metric of system scalability that does not depend on access to audio corpora.

音频指纹生成模型评估方法可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。