arXiv:2409.07827cs.SDcs.CV2024-09被引 11

用画作情绪生成匹配音乐,让艺术跨感官传递情感。

Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings

  • 先分析画作情绪并生成描述文本,再转为音乐
  • 在自建数据集上实现音乐与情绪文本高度一致
  • 适合残障辅助、教育疗愈等多感官应用

人工智能在音乐与图像生成方面取得显著进展,但跨模态情感对齐仍具挑战。本文提出一种双阶段模型,通过情绪标注、图像描述生成和语言模型,将视觉艺术中的情感转化为音乐作品。为解决艺术与音乐数据对齐稀缺问题,构建了情感绘画音乐数据集(Emotion Painting Music Dataset),用于训练与评估。模型先将画作转换为情感描述文本,再生成对应音乐,实现高效学习且仅需少量数据。使用弗雷谢特音频距离(FAD)、总谐波失真(THD)、Inception Score(IS)及KL散度等指标评估性能,结合预训练的CLAP模型验证生成音乐与文本间的情感一致性。该工具可连接视觉艺术与音乐,提升视障人群感知体验,拓展教育与治疗中的多感官应用可能性。

原文摘要 · Abstract (English)

Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that resonates with the emotions depicted in visual arts, integrating emotion labeling, image captioning, and language models to transform visual inputs into musical compositions. Addressing the scarcity of aligned art and music data, we curated the Emotion Painting Music Dataset, pairing paintings with corresponding music for effective training and evaluation. Our dual-stage framework converts images to text descriptions of emotional content and then transforms these descriptions into music, facilitating efficient learning with minimal data. Performance is evaluated using metrics such as Fréchet Audio Distance (FAD), Total Harmonic Distortion (THD), Inception Score (IS), and KL divergence, with audio-emotion text similarity confirmed by the pre-trained CLAP model to demonstrate high alignment between generated music and text. This synthesis tool bridges visual art and music, enhancing accessibility for the visually impaired and opening avenues in educational and therapeutic applications by providing enriched multi-sensory experiences.

跨模态生成情感计算音乐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。