用大模型分析用户音乐偏好并生成精准音乐生成提示
TuneGenie: Reasoning-based LLM agents for preferential music generation
- 基于用户歌单和文字描述,用大模型分析音乐偏好
- 生成的提示能有效提升Suno AI的音乐生成质量
- 适合对个性化音乐生成感兴趣的AI研究者
近期,大语言模型(LLMs)在图像生成、空间推理等多样化任务中展现出巨大潜力。鉴于其日益增强的文本推理能力,我们探究了LLMs在分析个人音乐偏好(基于歌单元数据、个人文字记录等)方面的表现,并据此生成有效提示,传递给Suno AI(一种音乐生成工具)。本文提出一种新型基于大模型的音乐文本表示方法(命名为TuneGenie),并开发了多种评估与基准测试方法,丰富了当前关于人工智能创作艺术的研究体系,该领域正日益增长且充满争议。
原文摘要 · Abstract (English)
Recently, Large language models (LLMs) have shown great promise across a diversity of tasks, ranging from generating images to reasoning spatially. Considering their remarkable (and growing) textual reasoning capabilities, we investigate LLMs' potency in conducting analyses of an individual's preferences in music (based on playlist metadata, personal write-ups, etc.) and producing effective prompts (based on these analyses) to be passed to Suno AI (a generative AI tool for music production). Our proposition of a novel LLM-based textual representation to music model (which we call TuneGenie) and the various methods we develop to evaluate & benchmark similar models add to the increasing (and increasingly controversial) corpus of research on the use of AI in generating art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。