用视觉纹理控制多轨音乐生成,让作曲者像画画一样操作乐器音色与布局。
ViTex: Visual Texture Control for Multi-Track Symbolic Music Generation via Discrete Diffusion Models
- 用颜色、位置、笔触构建乐器纹理的视觉表示,实现直观控制。
- 基于视觉纹理和和弦生成8小节多轨乐谱,质量媲美无条件生成模型。
- 适合音乐创作人、作曲家,想直接操控乐器搭配与声音质感的人使用。
在自动音乐生成中,如何设计有意义的人机交互控制是核心挑战。现有系统多依赖文本提示或元数据等外部输入,无法让用户直接干预作品结构。尽管已有研究尝试使用和弦或层次结构等内在控制,但主要集中在钢琴或声乐伴奏场景,对多轨符号化音乐的探索仍不足。本文识别出‘配器’(instrumentation)——即乐器选择及其角色分配——是多轨创作中的自然控制维度,提出 ViTex:一种乐器纹理的视觉表征方法。在 ViTex 中,颜色代表乐器类型,空间位置表示音高与时间,笔触属性刻画局部纹理特征。基于此表示,我们构建了一个以 ViTex 和和弦进行为条件的离散扩散模型,可生成8小节的多轨符号化音乐,在保持强大无条件生成质量的同时,实现精细的纹理级控制。演示页与代码已公开于 https://vitex2025.github.io/。
原文摘要 · Abstract (English)
In automatic music generation, a central challenge is to design controls that enable meaningful human-machine interaction. Existing systems often rely on extrinsic inputs such as text prompts or metadata, which do not allow humans to directly shape the composition. While prior work has explored intrinsic controls such as chords or hierarchical structure, these approaches mainly address piano or vocal-accompaniment settings, leaving multitrack symbolic music largely underexplored. We identify instrumentation, the choice of instruments and their roles, as a natural dimension of control in multi-track composition, and propose ViTex, a visual representation of instrumental texture. In ViTex, color encodes instrument choice, spatial position represents pitch and time, and stroke properties capture local textures. Building on this representation, we develop a discrete diffusion model conditioned on ViTex and chord progressions to generate 8-measure multi-track symbolic music, enabling explicit texture-level control while maintaining strong unconditional generation quality. The demo page and code are avaliable at https://vitex2025.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。