arXiv:2608.03433cs.HCcs.SD2026-08

跨文化测试AI生成音乐的味觉对应关系,发现感知差异主要源于评分习惯而非真实感知不同。

Cross-cultural evaluation of taste-sound correspondences in AI-generated music

论文配图:Cross-cultural evaluation of taste-sound correspondences in AI-generated music
图 1 · 摘自论文原文
  • 在阿根廷、意大利、日本开展三国在线实验,比较AI音乐对甜酸苦咸的味觉映射效果。
  • 多数国家偏好微调后的音乐,但日本组无显著偏好,咸味映射最弱。
  • 差异主要来自评分尺度使用习惯,但味觉-声音关联结构仍具跨文化一致性。

声学调味研究显示听众会系统性地将味觉与情绪意义赋予声音,而文本转音乐的生成式AI近期被用于将味觉提示转化为音乐刺激。这类模型所获得的味觉-声音对应关系是否超越其验证的文化背景尚不清楚。本研究将单国研究扩展至阿根廷、意大利和日本的三国家在线实验(N = 361)。参与者首先在基础版与微调版MusicGen音乐片段中选择偏好的版本,这些片段由四种味觉提示(甜、酸、苦、咸)生成;随后对微调版本在十二项味觉、情绪和温度描述词上进行评分。结果显示,在阿根廷与意大利,微调模型获得更高偏好,但在日本未达显著水平;所有三组中,咸味提示的对应关系最弱。各国评分差异显著,但若按个体内部标准化后,国家主效应不再显著,而提示与描述词之间的交互模式基本保持不变。因此,表面的跨文化差异主要源于评分尺度使用差异;结构性映射关系仍存在。此外探索性因子分析表明,各组的十二个描述词在潜维度上组织方式不同。结果表明,跨文化声学调味效应存在于两个层面:一是对特定刺激赋予味觉的整体程度,二是这些赋义之间的关系结构。评估生成式音乐系统在不同人群中的表现时,应区分反应风格偏差与真实的感知重构。

原文摘要 · Abstract (English)

Sonic seasoning research has shown that listeners attribute systematic gustatory and emotional meaning to sound, and text-to-music generative artificial intelligence has recently been used to render gustatory prompts as musical stimuli. Whether the taste-sound correspondences acquired by such models hold beyond the cultural context in which they were validated remains untested. We extended a single-country study to a three-country online experiment conducted in Argentina, Italy, and Japan (N = 361). Participants first indicated their preference between base and fine-tuned MusicGen excerpts generated from four taste prompts (sweet, sour, bitter, salty), and then rated fine-tuned excerpts on twelve taste, emotion, and thermal descriptors. Preference for the fine-tuned model was confirmed in Argentina and Italy but not in Japan, and the salty prompt yielded the weakest correspondence in all three cohorts. Ratings differed substantially between countries, yet the main effect of country was no longer detectable once ratings had been standardized within participant, whereas the interactions characterizing the mapping of prompts onto descriptors remained essentially unchanged. Much of the apparent cross-cultural divergence is therefore attributable to differences in scale use; a structural component nevertheless persists. In addition an exploratory factor analysis indicated that the twelve descriptors were organized along different latent dimensions in each cohort. These results indicate that cross-cultural variation in AI-mediated sonic seasoning operates at two levels: the overall level at which taste is attributed to a given stimulus, and the relational structure of those attributions. Evaluations of generative music systems across populations should accordingly distinguish response-style bias from genuine perceptual reorganization.

跨文化音乐生成声学调味AI感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。