用图像和评论生成情感匹配的音乐,无需人工标注情绪标签。
Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
- 通过融合图像与文本特征,生成与艺术作品情感一致的音乐。
- 在ArtiCaps数据集上,音高保真度与感知自然性显著提升。
- 适合交互艺术、个性化声音景观等创意场景使用。
随着AI生成内容(AIGC)的发展,从多模态输入生成感知自然且情感一致的音乐成为核心挑战。现有方法常依赖昂贵标注的情绪标签,亟需更灵活的情感对齐方案。为此,我们构建了ArtiCaps——一个通过语义匹配ArtEmis与MusicCaps描述生成的伪情感对齐图像-音乐-文本数据集。进一步提出Art2Music,一种轻量级跨模态框架,可从艺术图像和用户评论中合成音乐。第一阶段使用OpenCLIP编码图像与文本,并通过门控残差模块融合;融合表示经双向LSTM解码为梅尔频谱图,采用频率加权L1损失提升高频保真度。第二阶段通过微调的HiFi-GAN声码器重建高质量音频波形。在ArtiCaps上的实验表明,梅尔倒谱失真、Frechet音频距离、对数谱距离及余弦相似度均有显著改善。基于小型大模型的评分研究验证了跨模态情感一致性,并提供匹配与不匹配的可解释分析。结果表明,该方法在感知自然性、频谱保真度与语义一致性方面均表现优越。Art2Music仅需5万训练样本即可保持稳健性能,为互动艺术、个性化声景与数字艺术展览中的情感对齐音频生成提供了可扩展方案。
原文摘要 · Abstract (English)
With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches often rely on explicit emotion labels that require costly annotation, underscoring the need for more flexible feeling-aligned methods. To support multimodal music generation, we construct ArtiCaps, a pseudo feeling-aligned image-music-text dataset created by semantically matching descriptions from ArtEmis and MusicCaps. We further propose Art2Music, a lightweight cross-modal framework that synthesizes music from artistic images and user comments. In the first stage, images and text are encoded with OpenCLIP and fused using a gated residual module; the fused representation is decoded by a bidirectional LSTM into Mel-spectrograms with a frequency-weighted L1 loss to enhance high-frequency fidelity. In the second stage, a fine-tuned HiFi-GAN vocoder reconstructs high-quality audio waveforms. Experiments on ArtiCaps show clear improvements in Mel-Cepstral Distortion, Frechet Audio Distance, Log-Spectral Distance, and cosine similarity. A small LLM-based rating study further verifies consistent cross-modal feeling alignment and offers interpretable explanations of matches and mismatches across modalities. These results demonstrate improved perceptual naturalness, spectral fidelity, and semantic consistency. Art2Music also maintains robust performance with only 50k training samples, providing a scalable solution for feeling-aligned creative audio generation in interactive art, personalized soundscapes, and digital art exhibitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。