arXiv:2607.10023cs.SDcs.LG2026-07

用全局监督实现音乐与乐谱的局部对齐,提升跨模态匹配精度。

Local Multimodal Music Alignment from Global Supervision

论文配图:Local Multimodal Music Alignment from Global Supervision
图 1 · 摘自论文原文
  • 通过基于Sinkhorn的软对齐,直接计算图像块与音频帧的局部相似度
  • 仅需全局标注(如音视频片段配对)即可学习精细时间对齐关系
  • 适用于需要精确音乐时序对齐的任务,如乐谱-音频同步

理解音乐需把握多模态间的局部对应关系,例如演奏音频中的时间点如何对应乐谱图像中的位置。但这类局部标注难以获取,实际中常仅有音视频段落级别的全局标注。为此,我们提出FuSiLi(Fused Sinkhorn-Localized Similarity),一种在局部图像块与音频帧特征间直接计算相似度的对比学习方法,基于Sinkhorn实现软对齐。实验表明,FuSiLi(i)能有效学习局部关系,(ii)仅需全局监督,(iii)保留传统对比方法的全局对齐能力。我们在原始乐谱图像与音频对上,使用融合FuSiLi与传统全局相似性的混合对比目标微调预训练的CLIP与CLAP编码器。在跨模态检索与帧级对齐任务上评估,结果表明该方法在局部对齐上优于多种全局与局部基线,同时在检索任务上保持竞争力。

原文摘要 · Abstract (English)

Understanding music requires understanding localized relationships across data modalities, e.g., how time in performance audio maps onto position in a score image. Yet supervision for such local correspondences is difficult to obtain-in practice, we often only have access to coarser global supervision like paired segments of audio and images. To address this gap, we propose FuSiLi (Fused Sinkhorn-Localized Similarity), a similarity score for multimodal contrastive learning operating directly on local image patch and audio frame features via Sinkhorn-based soft alignment. We show that FuSiLi (i) effectively learns local relationships, (ii) requires only global supervision, and (iii) retains the global alignment capabilities of conventional contrastive approaches. We fine-tune pretrained CLIP and CLAP encoders on pairs of raw sheet music images and audio using a hybrid contrastive objective combining FuSiLi with conventional global similarity. We evaluate on cross-modal retrieval and frame-level alignment tasks against a range of global and local baselines, showing that our approach outperforms them on local alignment while remaining competitive on retrieval.

多模态对齐音乐理解对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。