arXiv:2608.06634cs.SD2026-08中稿 · the 27th Internati…

研究用户写音乐提示与描述听觉体验的差异,发现英文韩文用户表达方式不同。

From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music

论文配图:From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music
图 1 · 摘自论文原文
  • 分析200个真实音乐提示与听众描述,构建基于实际数据的音乐描述词库。
  • 提示中以流派和叙事语言为主,叙事类提示易导致理解偏差。
  • 跨文化对比显示英韩用户在情感、功能、叙事维度上描述习惯不同。

文本生成音乐(TTM)系统允许用户通过自然语言提示创作音乐,但尚不清楚提示用语是否与听觉体验的描述语言一致。本研究将200个真实世界中的Udio提示与其生成音频及英语(n=70)和韩语(n=78)听众的自由描述配对,贡献一个基于真实用户数据的人类主导的音乐提示词汇分类体系。结合词级与向量级分析,发现结构不对称:提示以流派和故事/叙事语言为主。流派术语在从提示到感知的传递中最可靠,而叙事性提示是语义错位最强的预测因子。初步跨文化比较进一步表明,不同听众群体在叙事、功能和情感维度上的描述模式存在差异,引发疑问:当前基于聚合英文语料训练的TTM系统能否容纳人们自然表达音乐想法的全部多样性。

原文摘要 · Abstract (English)

Text-to-music (TTM) generation systems allow users to create music through natural language prompts, yet it is unclear whether the descriptive language used to prompt aligns with descriptive language used to summarize or describe heard music. We pair 200 real-world Udio prompts with their generated audio and free-form descriptions collected from English- (n = 70) and Korean-speaking (n = 78) listeners, and contribute a human-derived taxonomy of musical prompting vocabulary grounded in real user data. Using this framework, alongside word- and vector-level analyses, we find a consistent structural asymmetry: prompts are dominated by Genre and Story/Narrative language. Genre terms propagate most reliably from prompt to perception, while narrative-heavy prompts are the strongest predictor of semantic misalignment. A preliminary cross-cultural comparison further suggests that description profiles vary across listener populations along narrative, functional, and affective dimensions, raising questions about whether current TTM systems, trained on aggregated English-centric corpora, can accommodate the full diversity of how people naturally express musical ideas.

文本生成音乐跨文化研究语言表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。