arXiv:2503.00427cs.SDcs.AI2025-03被引 2

提出跨模态语言模型映射,用音乐研究实现更高效的知识迁移。

Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal

  • 通过语言模型映射实现跨模态深层对齐,突破传统嵌入对齐局限。
  • 音乐领域数据稀缺,该方法可提升样本效率,减少对配对数据依赖。
  • 适合研究多模态学习、知识迁移及符号与感官融合的学者。

深度神经网络在表征学习和语言模型(LM)方面取得了显著进展。许多研究试图通过令牌或嵌入层面的对齐建立不同模态间的联系,但现有方法大多依赖大量数据,在音乐等配对数据稀缺的领域表现受限。我们认为,嵌入对齐仅停留在多模态对齐的表面层次。本文提出一项宏大挑战:语言模型映射(LMM),即在假设不同模态的语言模型追踪相同潜在现象的前提下,如何将一个领域语言模型中的本质信息映射到另一个领域。我们首先介绍LMM的基本框架,强调其揭示跨模态深层对齐并实现更高效的样本学习目标。随后讨论音乐作为开展LMM研究的理想领域的原因。接着,将音乐中的LMM与更普遍且更具挑战性的科学问题——基于感官输入和抽象符号进行决策——相联系,并最终呈现挑战问题的进阶版本设定。

原文摘要 · Abstract (English)

We have seen remarkable success in representation learning and language models (LMs) using deep neural networks. Many studies aim to build the underlying connections among different modalities via the alignment and mappings at the token or embedding level, but so far, most methods are very data-hungry, limiting their performance in domains such as music where paired data are less abundant. We argue that the embedding alignment is only at the surface level of multimodal alignment. In this paper, we propose a grand challenge of \textit{language model mapping} (LMM), i.e., how to map the essence implied in the LM of one domain to the LM of another domain under the assumption that LMs of different modalities are tracking the same underlying phenomena. We first introduce a basic setup of LMM, highlighting the goal to unveil a deeper aspect of cross-modal alignment as well as to achieve more sample-efficiency learning. We then discuss why music is an ideal domain in which to conduct LMM research. After that, we connect LMM in music with a more general and challenging scientific problem of \textit{learning to take actions based on both sensory input and abstract symbols}, and in the end, present an advanced version of the challenge problem setup.

多模态学习语言模型音乐生成知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。