arXiv:2504.13891cs.HCcs.AI2025-04被引 4

用文字、图片等生成情感化音乐,支持多风格融合创作。

Mozualization: Crafting Music and Visual Representation with Multimodal AI

  • 输入关键词/图像/音效,生成融合多风格的音乐
  • 用户研究显示9人参与,普遍认可其情感表达能力
  • 适合音乐创作者与情感表达爱好者使用

本文提出Mozualization,一款音乐生成与编辑工具,通过整合关键词、图像和声音片段(如不同乐曲段落或猫咪叫声)等多元输入,生成多风格融合的嵌入式音乐。该工作受人类情感表达方式启发——如写情绪诗文、绘制冷暖色调画作或聆听悲伤/振奋音乐。基于此,我们构建了一个将情感表达转化为连贯且富有表现力歌曲的工具,使用户能无缝融入个人偏好与灵感。为评估工具效果并获取优化反馈,我们开展了包含九位音乐爱好者的用户研究,重点考察用户体验、参与度以及与生成音乐互动后的感受。

原文摘要 · Abstract (English)

In this work, we introduce Mozualization, a music generation and editing tool that creates multi-style embedded music by integrating diverse inputs, such as keywords, images, and sound clips (e.g., segments from various pieces of music or even a playful cat's meow). Our work is inspired by the ways people express their emotions -- writing mood-descriptive poems or articles, creating drawings with warm or cool tones, or listening to sad or uplifting music. Building on this concept, we developed a tool that transforms these emotional expressions into a cohesive and expressive song, allowing users to seamlessly incorporate their unique preferences and inspirations. To evaluate the tool and, more importantly, gather insights for its improvement, we conducted a user study involving nine music enthusiasts. The study assessed user experience, engagement, and the impact of interacting with and listening to the generated music.

音乐生成多模态情感表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。