arXiv:2604.10283cs.SDcs.LG2026-04

用手工特征增强音频与乐谱对齐,提升跨模态检索效果

Descriptor-Injected Cross-Modal Learning: A Systematic Exploration of Audio-MIDI Alignment via Spectral and Melodic Features

  • 在编码器中注入频带能量等手工特征作为桥梁
  • 最佳模型在五次随机实验中平均准确率达84.0%,提升8.8个百分点
  • 高频音阶带是关键判别信号,适合音乐信息检索研究者

音频与符号化乐谱(MIDI)之间的跨模态检索仍具挑战,因连续波形与离散事件序列分别表征表演的不同方面。本文研究描述符注入——在模态专用编码器中引入手工设计的领域特征,以弥合这一差距。通过涵盖13种描述符-机制组合、6类架构和3种训练策略的三阶段实验,最优配置在五次独立种子测试中达到84.0%的平均准确率,相比无描述符基线提升8.8个百分点。因果消融显示,基于八度带能量动态的音频描述符A4驱动了顶级双模模型的性能提升,而MIDI描述符D4虽改善训练动态,但在推理阶段影响微弱。我们还提出反向交叉注意力机制,使描述符令牌主动查询编码特征,在减少注意力计算的同时保持竞争力。核中心化相关分析(CKA)表明,描述符显著增强了音频-MIDI Transformer层间的表征对齐,说明其促进的是表征收敛而非简单拼接。扰动分析识别出高频八度带为最主要的判别信号。所有实验均基于MAESTRO v3.0.0数据集,评估协议控制作曲家与曲目相似性。

原文摘要 · Abstract (English)

Cross-modal retrieval between audio recordings and symbolic music representations (MIDI) remains challenging because continuous waveforms and discrete event sequences encode different aspects of the same performance. We study descriptor injection, the augmentation of modality-specific encoders with hand-crafted domain features, as a bridge across this gap. In a three-phase campaign covering 13 descriptor-mechanism combinations, 6 architectural families, and 3 training schedules, the best configuration reaches a mean S of 84.0 percent across five independent seeds, improving the descriptor-free baseline by 8.8 percentage points. Causal ablation shows that the audio descriptor A4, based on octave-band energy dynamics, drives the gain in the top dual models, while the MIDI descriptor D4 has only a weak inference-time effect despite improving training dynamics. We also introduce reverse cross-attention, where descriptor tokens query encoder features, reducing attention operations relative to the standard formulation while remaining competitive. CKA analysis shows that descriptors substantially increase audio-MIDI transformer layer alignment, indicating representational convergence rather than simple feature concatenation. Perturbation analysis identifies high-frequency octave bands as the dominant discriminative signal. All experiments use MAESTRO v3.0.0 with an evaluation protocol controlling for composer and piece similarity.

跨模态音乐生成特征注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。