用大模型分析混响等音效如何影响情绪,揭示声音设计的感知机制。
Exploring How Audio Effects Alter Emotion with Foundation Models
- 用预训练大模型分析音效对情绪的影响
- 发现特定音效与情绪变化存在非线性关联
- 适合音乐认知与情感计算研究者参考
混响、失真、调制和动态范围处理等音频效果在音乐聆听中对情绪反应起关键作用。尽管以往研究关注低层音频特征与情感感知的关系,但音频效果对情绪的系统性影响仍待深入探索。本文利用基础模型——在多模态数据上预训练的大规模神经架构——分析这些效果。此类模型编码了音乐结构、音色与情感意义之间的丰富关联,为探究声音设计技术的情绪后果提供了强大框架。通过多种探测方法分析深度学习模型的嵌入表示,我们考察了音频效果与估计情绪间的复杂非线性关系,揭示了特定效果对应的情感模式,并评估了基础音频模型的鲁棒性。研究结果有助于深化对音频制作实践感知影响的理解,对音乐认知、表演及情感计算具有重要意义。
原文摘要 · Abstract (English)
Audio effects (FX) such as reverberation, distortion, modulation, and dynamic range processing play a pivotal role in shaping emotional responses during music listening. While prior studies have examined links between low-level audio features and affective perception, the systematic impact of audio FX on emotion remains underexplored. This work investigates how foundation models - large-scale neural architectures pretrained on multimodal data - can be leveraged to analyze these effects. Such models encode rich associations between musical structure, timbre, and affective meaning, offering a powerful framework for probing the emotional consequences of sound design techniques. By applying various probing methods to embeddings from deep learning models, we examine the complex, nonlinear relationships between audio FX and estimated emotion, uncovering patterns tied to specific effects and evaluating the robustness of foundation audio models. Our findings aim to advance understanding of the perceptual impact of audio production practices, with implications for music cognition, performance, and affective computing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。