arXiv:2605.22255cs.CVcs.IR2026-05被引 1

让音乐谱图直接支持内容搜索,突破传统元数据限制。

Direct content-based retrieval from music scores images

论文配图:Direct content-based retrieval from music scores images
图 1 · 摘自论文原文
  • 构建可复用的查询数据集生成方法,聚焦关键谱图特征。
  • 多模型对比:OMR管道在同域检索更准,无转录模型抗域变化更强。
  • 适合音乐学者、教育者及数字典藏项目快速定位谱例。

乐谱数字化对保存与获取至关重要,但信息检索仍依赖标题或作曲家等元数据,基于内容的乐谱图像检索远未充分探索。本文首先研究乐谱中对搜索最相关的特征,并提出从任意标注语料库系统构建查询数据集的方法。对比多种内容检索方案:基于光学乐谱识别(OMR)的转录方法、直接从谱图识别查询的无转录Transformer模型,以及文本提示的大语言模型。在四个具有不同规模、图像质量与排版方式的语料库上评估,结果表明:OMR方法在同域场景下表现更优,而无转录模型更能适应领域差异。

原文摘要 · Abstract (English)

The digitization of musical scores plays a crucial role in their preservation and accessibility, yet information retrieval still depends mainly on metadata searches, such as by title or composer. Content based search in music score images remains underexplored compared to text documents, despite its potential value for musicians, musicologists, and educators. This work contributes to the field by first studying which characteristics of a score are most relevant for search and by defining a systematic method to build query datasets from any annotated corpus. We also consider diverse methods for content-based search on music score images, ranging from transcription-based approaches relying on Optical Music Recognition (OMR), to a transcription-free Transformer model trained to recognize queries directly from score images, and a text-prompted Large Language Model. Our experiments evaluate these models on four corpora exhibiting diverse characteristics in terms of dataset size, image quality, and typesetting mechanisms. Overall, each method excels under different conditions: OMR-based pipelines achieve higher in-domain retrieval, whereas transcription-free models handle domain variability more effectively.

音乐信息检索光学乐谱识别图像搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。