验证音乐与风味的跨模态关联在合成数据中依然有效,推动可复现的多模态研究。
Multimodal Dataset Normalization and Perceptual Validation for Music-Taste Correspondences
- 通过实验验证音频风味关联在不同数据集间可迁移
- 计算生成的风味目标与人类感知高度一致(相关性r=0.45)
- 适合多模态AI、音乐感知与食品风味交叉研究者
构建大规模对齐的跨模态数据集在音乐风味研究中面临挑战,因感知实验成本高且样本量小。本文通过两个互补实验解决此瓶颈:第一个实验检验来自257首带人工标注的音轨集合中的音频-风味关联、特征重要性排序及潜在因子结构,能否迁移到约49,300个片段的合成标签FMA语料库;第二个实验在在线听觉研究中(49名参与者,20首曲目),验证基于食品化学的可复现流程生成的计算风味目标是否与人类感知一致。结果表明:定量迁移分析确认跨监督模式下的跨模态结构保持稳定;感知评估显示计算目标与听者评分显著一致(置换检验p<0.0001,Mantel r=0.45,Procrustes m²=0.51)。两项结果共同支持结论:合成FMA标注中存在声学调味效应。论文发布数据集和配套代码,以支持可复现的跨模态人工智能研究。
原文摘要 · Abstract (English)
Collecting large, aligned cross-modal datasets for music-flavor research is difficult because perceptual experiments are costly and small by design. We address this bottleneck through two complementary experiments. The first tests whether audio-flavor correlations, feature-importance rankings, and latent-factor structure transfer from an experimental soundtracks collection (257~tracks with human annotations) to a large FMA-derived corpus ($\sim$49,300 segments with synthetic labels). The second validates computational flavor targets -- derived from food chemistry via a reproducible pipeline -- against human perception in an online listener study (49~participants, 20~tracks). Results from both experiments converge: the quantitative transfer analysis confirms that cross-modal structure is preserved across supervision regimes, and the perceptual evaluation shows significant alignment between computational targets and listener ratings (permutation $p<0.0001$, Mantel $r=0.45$, Procrustes $m^2=0.51$). Together, these findings support the conclusion that sonic seasoning effects are present in synthetic FMA annotations. We release datasets and companion code to support reproducible cross-modal AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。