arXiv:2412.11818cs.MMcs.IR2024-12中稿 · presentation at NL…被引 1

利用用户生成的视频元数据提升YouTube翻唱歌曲识别准确率

Leveraging User-Generated Metadata of Online Videos for Cover Song Identification

  • 融合实体消歧模型与音频分析,构建多模态识别框架
  • 引入用户元数据后,识别性能在多个数据集上显著提升
  • 适合对在线音乐内容分析感兴趣的开发者与研究者

YouTube 是翻唱歌曲的重要来源。由于平台以视频为组织单位而非歌曲,翻唱识别难度较大。传统方法依赖音频内容,而用户生成的视频元数据有望提升识别效果。本文提出一种多模态翻唱识别方法,结合实体消歧模型与基于音频的方法,通过排序模型进行融合。实验表明,利用用户生成的元数据可有效稳定 YouTube 上的翻唱识别性能。

原文摘要 · Abstract (English)

YouTube is a rich source of cover songs. Since the platform itself is organized in terms of videos rather than songs, the retrieval of covers is not trivial. The field of cover song identification addresses this problem and provides approaches that usually rely on audio content. However, including the user-generated video metadata available on YouTube promises improved identification results. In this paper, we propose a multi-modal approach for cover song identification on online video platforms. We combine the entity resolution models with audio-based approaches using a ranking model. Our findings implicate that leveraging user-generated metadata can stabilize cover song identification performance on YouTube.

翻唱识别多模态元数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。