arXiv:2509.22378cs.SDcs.AI2025-09

用视觉语言模型实现无需训练的可解释图像转音乐生成

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach

  • 基于VLM和ABC记谱法,通过自然语言生成音乐
  • 零训练成本下达成高质量音乐与图像一致性的生成结果
  • 支持文本和图像双模态解释,适合艺术创作与教育场景

图像到音乐(I2M)生成近年来受到广泛关注,潜在应用涵盖游戏、广告和多模态艺术创作。然而,由于I2M任务本身具有模糊性和主观性,多数端到端方法缺乏可解释性,用户难以理解生成结果。即便基于情绪映射的方法也存在争议,因情绪仅是艺术的一个维度。此外,大多数学习型方法需要大量计算资源和大规模数据集训练,限制了普通用户的使用。为此,我们提出首个基于视觉语言模型(VLM)的I2M框架,兼具高可解释性与低计算开销。具体而言,采用ABC记谱法连接文本与音乐模态,使VLM可通过自然语言生成音乐;结合多模态检索增强生成(RAG)与自精炼技术,实现无需外部训练的高质量音乐生成;同时利用生成的文本动机与VLM注意力图,提供跨模态的解释。通过人评与机器评估验证,本方法在音乐质量与音乐-图像一致性上均优于现有方法,表现优异。代码已开源。

原文摘要 · Abstract (English)

Recently, Image-to-Music (I2M) generation has garnered significant attention, with potential applications in fields such as gaming, advertising, and multi-modal art creation. However, due to the ambiguous and subjective nature of I2M tasks, most end-to-end methods lack interpretability, leaving users puzzled about the generation results. Even methods based on emotion mapping face controversy, as emotion represents only a singular aspect of art. Additionally, most learning-based methods require substantial computational resources and large datasets for training, hindering accessibility for common users. To address these challenges, we propose the first Vision Language Model (VLM)-based I2M framework that offers high interpretability and low computational cost. Specifically, we utilize ABC notation to bridge the text and music modalities, enabling the VLM to generate music using natural language. We then apply multi-modal Retrieval-Augmented Generation (RAG) and self-refinement techniques to allow the VLM to produce high-quality music without external training. Furthermore, we leverage the generated motivations in text and the attention maps from the VLM to provide explanations for the generated results in both text and image modalities. To validate our method, we conduct both human studies and machine evaluations, where our method outperforms others in terms of music quality and music-image consistency, indicating promising results. Our code is available at https://github.com/RS2002/Image2Music .

图像转音乐可解释生成视觉语言模型ABC记谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。