用画作直接生成音乐,跳过文字中间环节。
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
- 直接从画作映射到音乐,不依赖文字描述。
- 基于10.6万对画作-音乐数据训练,保持视觉与音乐风格一致。
- 适合多媒体艺术、文化遗产数字化等创意应用。
音乐生成在多模态深度学习推动下取得显著进展,可从文本甚至图像生成音频。但现有图像条件系统存在两大局限:(i)通常在自然照片上训练,难以捕捉艺术品丰富的语义、风格与文化内涵;(ii)多数依赖图像转文字阶段,以语言作为语义捷径,简化了条件输入却阻碍了直接的视觉到音频学习。为此,我们构建了ArtSound,一个包含105,884对画作-音乐样本的大规模多模态数据集,附带双模态标题,由ArtGraph和Free Music Archive扩展而来。同时提出ArtToMus,首个专为直接画作到音乐生成设计的框架,无需图像转文字或语言语义监督。该框架将视觉嵌入投影至潜在扩散模型的条件空间,实现仅基于视觉信息的音乐合成。实验表明,ArtToMus生成的音乐在旋律连贯性与风格一致性上表现良好,能反映源画作的关键视觉特征。尽管绝对对齐分数低于文本条件系统——这在去除语言监督后属预期结果——但其感知质量具有竞争力,且具备有意义的跨模态对应关系。本工作确立了直接视觉到音乐生成这一独特而具挑战性的研究方向,并提供了支持多媒体艺术、文化遗产及AI辅助创作的应用资源。代码与数据集将在论文接收后公开。
原文摘要 · Abstract (English)
Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-conditioned systems suffer from two fundamental limitations: (i) they are typically trained on natural photographs, limiting their ability to capture the richer semantic, stylistic, and cultural content of artworks; and (ii) most rely on an image-to-text conversion stage, using language as a semantic shortcut that simplifies conditioning but prevents direct visual-to-audio learning. Motivated by these gaps, we introduce ArtSound, a large-scale multimodal dataset of 105,884 artwork-music pairs enriched with dual-modality captions, obtained by extending ArtGraph and the Free Music Archive. We further propose ArtToMus, the first framework explicitly designed for direct artwork-to-music generation, which maps digitized artworks to music without image-to-text translation or language-based semantic supervision. The framework projects visual embeddings into the conditioning space of a latent diffusion model, enabling music synthesis guided solely by visual information. Experimental results show that ArtToMus generates musically coherent and stylistically consistent outputs that reflect salient visual cues of the source artworks. While absolute alignment scores remain lower than those of text-conditioned systems-as expected given the substantially increased difficulty of removing linguistic supervision-ArtToMus achieves competitive perceptual quality and meaningful cross-modal correspondence. This work establishes direct visual-to-music generation as a distinct and challenging research direction, and provides resources that support applications in multimedia art, cultural heritage, and AI-assisted creative practice. Code and dataset will be publicly released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。