让代码生成更懂音乐,通过映射代码与音频嵌入提升多样性。
Embedding Alignment in Code Generation for Audio
- 构建代码与音频嵌入的非线性映射关系,实现双向对齐。
- 模型可预测代码生成的音频嵌入,提升输出音乐多样性。
- 适合音乐创作、生成式编程及跨模态交互研究者。
基于大语言模型的代码生成有望革新创意编码领域,如现场编程,使用户聚焦于结构设计而非语法细节。在该场景中,提供多样化的代码候选有助于更好地实现音乐意图。然而,现有代码生成模型难以产出独特且多样的代码,且缺乏对生成音频输出的直接理解。为此,我们研究了代码与音频嵌入空间之间的拓扑关系,发现二者不存在简单的线性对应,但通过构建预测模型表明可学习一个嵌入对齐映射。在此基础上,提出一种新模型,能根据输入代码预测对应的音频嵌入,从而建立代码-音频嵌入对齐图谱,以促进音乐风格多样化输出。
原文摘要 · Abstract (English)
LLM-powered code generation has the potential to revolutionize creative coding endeavors, such as live-coding, by enabling users to focus on structural motifs over syntactic details. In such domains, when prompting an LLM, users may benefit from considering multiple varied code candidates to better realize their musical intentions. Code generation models, however, struggle to present unique and diverse code candidates, with no direct insight into the code's audio output. To better establish a relationship between code candidates and produced audio, we investigate the topology of the mapping between code and audio embedding spaces. We find that code and audio embeddings do not exhibit a simple linear relationship, but supplement this with a constructed predictive model that shows an embedding alignment map could be learned. Supplementing the aim for musically diverse output, we present a model that given code predicts output audio embedding, constructing a code-audio embedding alignment map.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。