arXiv:2410.09928cs.SDcs.AI2024-10被引 4

用大模型自动为日漫生成贴合剧情的背景音乐。

M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models

  • 通过对话和角色表情识别场景边界与情绪,生成音乐指令。
  • 经主观评测,生成音乐更贴合剧情、质量更高且连贯性更强。
  • 适合对动漫音乐生成感兴趣的开发者与内容创作者。

本文提出M2M Gen,一个用于为日漫自动生成背景音乐的多模态框架。该任务面临缺乏数据集和基线的挑战。我们构建了自动化音乐生成流水线:首先利用漫画对话检测场景边界,并通过角色面部识别情绪;随后使用GPT4o将低层场景信息转化为高层音乐指令;再以该指令为条件,由另一实例GPT4o生成页面级音乐描述,指导文本到音乐模型生成与叙事同步的音乐。主观评估显示,相较基线方法,M2M Gen生成的音乐在相关性、一致性和质量上均有显著提升。

原文摘要 · Abstract (English)

This paper introduces M2M Gen, a multi modal framework for generating background music tailored to Japanese manga. The key challenges in this task are the lack of an available dataset or a baseline. To address these challenges, we propose an automated music generation pipeline that produces background music for an input manga book. Initially, we use the dialogues in a manga to detect scene boundaries and perform emotion classification using the characters faces within a scene. Then, we use GPT4o to translate this low level scene information into a high level music directive. Conditioned on the scene information and the music directive, another instance of GPT 4o generates page level music captions to guide a text to music model. This produces music that is aligned with the mangas evolving narrative. The effectiveness of M2M Gen is confirmed through extensive subjective evaluations, showcasing its capability to generate higher quality, more relevant and consistent music that complements specific scenes when compared to our baselines.

音乐生成多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。