arXiv:2509.00654cs.SDcs.AI2025-09

不用艺术家名字也能精准控制音乐风格,效果接近命名提示。

The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation

  • 用大模型生成可读描述词替代艺术家名,实现风格控制
  • 无名描述词恢复了约80%的命名提示控制力
  • 适合需合规、避免版权风险的音乐生成场景

文本生成音乐模型能捕捉乐器或情绪等宏观特征,但精细风格控制仍是难题。现有方法多需重新训练或特殊条件,影响复现性,且在禁止使用艺术家姓名时难以合规。本文研究是否可通过大语言模型生成的轻量级、人类可读描述词,提供政策友好的风格控制方案。以MusicGen-small为基础,评估比莉·艾利什(人声流行)与卢多维科·艾娜迪(钢琴器乐)两位艺术家,每人均使用15段参考片段,在三种条件下对比:基础提示、含艺术家名提示、五组描述词集。所有提示由大模型生成。采用VGGish与CLAP嵌入进行分布与逐片段相似度评估,引入新的最小距离归因指标。结果表明,艺术家名是最强控制信号,而无名描述词可恢复大部分效果。跨艺术家迁移导致对齐下降,说明描述词编码了特定风格线索。此外,构建了十位当代艺术家的描述词表。这些发现定义了‘无名差距’——即命名提示与合规描述词之间的可控性差异,并建立可复现的提示级控制评估协议。

原文摘要 · Abstract (English)

Text-to-music models capture broad attributes such as instrumentation or mood, but fine-grained stylistic control remains an open challenge. Existing stylization methods typically require retraining or specialized conditioning, which complicates reproducibility and limits policy compliance when artist names are restricted. We study whether lightweight, human-readable modifiers sampled from a large language model can provide a policy-robust alternative for stylistic control. Using MusicGen-small, we evaluate two artists: Billie Eilish (vocal pop) and Ludovico Einaudi (instrumental piano). For each artist, we use fifteen reference excerpts and evaluate matched seeds under three conditions: baseline prompts, artist-name prompts, and five descriptor sets. All prompts are generated using a large language model. Evaluation uses both VGGish and CLAP embeddings with distributional and per-clip similarity measures, including a new min-distance attribution metric. Results show that artist names are the strongest control signal across both artists, while name-free descriptors recover much of this effect. This highlights that existing safeguards such as the restriction of artist names in music generation prompts may not fully prevent style imitation. Cross-artist transfers reduce alignment, showing that descriptors encode targeted stylistic cues. We also present a descriptor table across ten contemporary artists to illustrate the breadth of the tokens. Together these findings define the name-free gap, the controllability difference between artist-name prompts and policy-compliant descriptors, shown through a reproducible evaluation protocol for prompt-level controllability.

音乐生成风格控制合规生成提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。