情绪音频建模的泛化失败并非数据不足,而是存在四种不可逾越的边界。
Not all generalisation failures can be bought back: four boundaries in affective audio modelling

- 通过跨数据集、音效编辑、生理信号替代自评等四类边界测试
- 跨音乐与环境声场景时性能损失达80%,预训练模型无法弥补
- 需区分泛化失败类型,错误应对会带来更高成本
将声学特征映射到情感反应的模型支撑音乐推荐与声音设计等应用,但其评估几乎仅限于训练数据集内部。当模型在外部失效时,常规做法是增加数据或扩大模型规模,假设所有失败都是资源不足所致。本文证明这并不正确,且应对方式应相反。我们使用四个标注声音数据集、四种预训练表征和三个生理信号数据集,测试模型在四个关键边界上的表现:新内容、音频编辑、传感器替代自评、个体听众差异。每项均报告目标上限、跨界存活率及缩小差距所需目标样本量。同一数据集内预测可达听众间一致性上限的84%;同域数据集切换损失约五分之一,100个目标标签可恢复三分之二损失;跨音乐与环境声场景则损失达五分之四,四种预训练表征均未恢复任何性能;针对生理反应,所构建的信息源均未超过理论上限的三分之一。因此,“模型不泛化”实为两种不同诊断,对应不同修复策略;将后者误作前者处理,代价更高。
原文摘要 · Abstract (English)
Models mapping acoustic properties onto affective response underpin applications from music recommendation to sound design, yet are evaluated almost entirely within the corpus they were fitted on. When one fails outside it, the standard response -- more data, or a larger model -- assumes every failure is a shortage of resources. We show it is not, and that the alternative calls for the opposite remedy. Using four corpora of rated sound, four pretrained representations and three corpora of physiological recording, we pushed one mapping across four boundaries an application must cross: to new material, to edited audio, to a sensor in place of a self-report, and to an individual listener. At each we report the ceiling the target permits, the fraction surviving the crossing, and the price in target-side observations of closing the gap. Within a corpus, prediction reaches 84% of the ceiling set by inter-listener agreement. A same-domain corpus swap costs a fifth of that, and a hundred target labels return two-thirds of the loss. Crossing between music and environmental sound costs four-fifths to all of it, and four pretrained representations recover none of it. Against physiological response no information source we constructed exceeds a third of the attainable ceiling. "The model does not generalise" is therefore two diagnoses, not one, with mutually exclusive remedies; treating the second as the first is the more expensive mistake.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。