用AI根据游戏进程动态生成配乐,让桌游音乐更贴合剧情。
Long-Form Text-to-Music Generation with Adaptive Prompts: A Case Study in Tabletop Role-Playing Games Soundtracks
- 用大模型将对话转为音乐描述,控制音乐生成。
- 详细音乐描述提升音质,连贯描述增强剧情契合度。
- 适合做桌游、游戏等需要长时序配乐的场景。
本文研究文本到音频音乐生成模型在生成随时间变化提示的长时序音乐方面的能力,聚焦于桌面角色扮演游戏(TRPG)配乐生成。我们提出Babel Bardo系统,利用大语言模型(LLMs)将语音转录文本转化为用于控制文本到音乐模型的音乐描述。在两个TRPG战役中对比了四种版本:一种基线使用直接语音转录,三种基于LLM的版本采用不同音乐描述生成方法。评估涵盖音频质量、剧情契合度和过渡流畅性。结果表明,详细的音乐描述可提升音频质量,而保持连续描述的一致性有助于增强剧情契合度与过渡平滑性。
原文摘要 · Abstract (English)
This paper investigates the capabilities of text-to-audio music generation models in producing long-form music with prompts that change over time, focusing on soundtrack generation for Tabletop Role-Playing Games (TRPGs). We introduce Babel Bardo, a system that uses Large Language Models (LLMs) to transform speech transcriptions into music descriptions for controlling a text-to-music model. Four versions of Babel Bardo were compared in two TRPG campaigns: a baseline using direct speech transcriptions, and three LLM-based versions with varying approaches to music description generation. Evaluations considered audio quality, story alignment, and transition smoothness. Results indicate that detailed music descriptions improve audio quality while maintaining consistency across consecutive descriptions enhances story alignment and transition smoothness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。