通过提示工程可精准复现特定音乐家的风格,揭示AI音乐生成中的版权隐忧。
The Artist is Present: Traces of Artists Resigind and Spawning in Text-to-Audio AI
- 用元标签设计提示词,定位并生成特定艺术家的音色特征。
- 成功复现博恩·伊弗、菲利普·格拉斯等音乐家的风格,且效果稳定。
- 为审计AI音乐生成的风格诱导性提供可复现方法,适合关注AI伦理者。
文本到音频(TTA)系统正迅速改变音乐创作与传播,平台如Udio和Suno每日生成数千首曲目并融入主流音乐生态。这些系统基于大规模且大多未公开的数据集训练,从根本上重塑了音乐的生产、复制与消费方式。本文通过元标签驱动的提示工程,实证发现可系统性微定位艺术家相关风格区域,实现对特定艺术家风格内容的“生成”。研究利用公开音乐分类体系中的描述符组合,成功重现了博恩·伊弗(Bon Iver)、菲利普·格拉斯(Philip Glass)、熊猫熊(Panda Bear)及威廉·巴辛斯基(William Basinski)等艺术家的典型音色特征。结果表明,文本-音频映射关系稳定,符合艺术家专属的训练信号,可在不明确提及艺术家姓名的情况下,精确导航至其风格微位置。该能力证明艺术家作品已成为系统生成新内容的基础素材,且常未经明确授权或署名。从概念上,揭示了文本描述符在高维表示空间中的导航作用;方法上,提供了可复现的风格可诱导性审计协议。研究引发关于治理、署名、同意与披露标准的紧迫问题,并挑战了创作权、复制、模仿、创作主体性与算法创作伦理的边界。
原文摘要 · Abstract (English)
Text-to-audio (TTA) systems are rapidly transforming music creation and distribution, with platforms like Udio and Suno generating thousands of tracks daily and integrating into mainstream music platforms and ecosystems. These systems, trained on vast and largely undisclosed datasets, are fundamentally reshaping how music is produced, reproduced and consumed. This paper presents empirical evidence that artist-conditioned regions can be systematically microlocated through metatag-based prompt design, effectively enabling the spawning of artist-like content through strategic prompt engineering. Through systematic exploration of metatag-based prompt engineering techniques this research reveals how users can access the distinctive sonic signatures of specific artists, evidencing their inclusion in training datasets. Using descriptor constellations drawn from public music taxonomies, the paper demonstrates reproducible proximity to artists such as Bon Iver, Philip Glass, Panda Bear and William Basinski. The results indicate stable text-audio correspondences consistent with artist-specific training signals, enabling precise traversal of stylistic microlocations without explicitly naming artists. This capacity to summon artist-specific outputs shows that artists' creative works fuction as foundational material from which these systems generate new content, often without explicit consent or attribuition. Conceptually, the work clarifies how textual descriptors act as navigational cues in high-dimensional representation spaces; methodologically, it provides a replicable protocol for auditing stylistic inducibility. The findings raise immediate queestions for governance-attribution, consent and disclosure standards-and for creative practice, where induced stylistic proximity complicates boundaries between ownership, reproduction, imitation, creative agency and the ethics of algorithmic creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。