提出可处理任意长度音频的指纹模型,提升真实场景识别准确率。
Variable-Length Audio Fingerprinting
- 采用可变长度输入,动态适应音频片段时长
- 在三个真实数据集上超越现有最优方法
- 适合实际应用中长短不一的音频识别
音频指纹技术将音频压缩为低维表示,使失真录音仍能被识别为原始内容。现有深度学习方法仅支持固定长度音频片段的指纹生成,忽略了分割过程中的时间动态性。为解决这一刚性限制,本文提出可变长度音频指纹(VLAFP),是首个支持训练与测试阶段任意长度音频输入的深度音频指纹模型。实验表明,该模型在三个真实世界数据集上的实时音频识别和音频检索任务中均优于现有最先进方法。
原文摘要 · Abstract (English)
Audio fingerprinting converts audio to much lower-dimensional representations, allowing distorted recordings to still be recognized as their originals through similar fingerprints. Existing deep learning approaches rigidly fingerprint fixed-length audio segments, thereby neglecting temporal dynamics during segmentation. To address limitations due to this rigidity, we propose Variable-Length Audio FingerPrinting (VLAFP), a novel method that supports variable-length fingerprinting. To the best of our knowledge, VLAFP is the first deep audio fingerprinting model capable of processing audio of variable length, for both training and testing. Our experiments show that VLAFP outperforms existing state-of-the-arts in live audio identification and audio retrieval across three real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。