用音频预测音乐带来的味觉感受,模型效果优于人类平均判断。
Taste-aware music retrieval from audio embeddings
- 用十种冻结的音频编码器+门控融合,从音乐音频预测五种味觉。
- 模型在未见音乐上误差仅0.13(人平均0.28),优于之前最佳基线0.219。
- 适合对跨模态感知、音乐情感计算感兴趣的研究者。
声音与味觉之间的跨模态对应关系在心理学和神经科学中已有研究,但在基于内容的多媒体检索中仍罕见。本文将音觉预测形式化为一个基于感知验证的多源语料库上的音乐信息检索基准,比较了来自四个HEAR家族的十种冻结音频编码器,在共享多任务回归头下,采用门控晚期融合作为可配置变体。为评估模型效果,计算绝对误差和秩相关性。最强系统在五种味觉预测上的宏均方根误差(RMSE)为0.134;在未见过的真实音乐上,其误差小于单个评价者的偏差(0.13对比0.28),表现优于平均人类评分者,显著低于此前最优基线(0.219)。在绝对误差上,各编码器统计无差异,单一VGGish即达最优融合效果;但门控晚期融合在秩相关性上优势明显(宏皮尔逊相关系数0.724对比0.666)。将预测味觉空间作为基于内容的检索索引,其对309项音乐池的排序精度远超文本基线CLAP(随机水平),且通过岭探针与音频带阻打击测试,验证了强表示与已知声-味关联的一致性。
原文摘要 · Abstract (English)
Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-source corpus, comparing ten frozen audio encoders from the four HEAR families under a shared multi-task regression head, with gated late-fusion as a configurable variant. In order to assess the effectiveness of the models, we compute absolute error and rank correlation. The strongest systems predict the five tastes within a macro RMSE of 0.134; on held-out real music their error is less than half a single rater's deviation from the consensus (RMSE 0.13 vs. 0.28), so the model tracks the group consensus more closely than an average human rater, and well below the previous state of the art baseline (0.219). On absolute error the encoders are statistically flat, with a single VGGish matching the best fusion, but gated late-fusion's advantage is confined to rank correlation (macro Pearson r 0.724 vs. 0.666). Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound-taste correspondences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。