揭示大脑如何分层编码声音情绪,高阶语义特征起主导作用。
Toward a Realistic Encoding Model of Auditory Affective Understanding in the Brain
- 基于听觉层级神经机制,分解声音为多层特征并建模情绪响应
- 高阶语义表征比低阶声学特征更能预测行为与神经同步(p<0.05)
- 人声在前额叶/颞叶激活更强,背景音在边缘系统更显著
在情感神经科学与情绪感知AI中,复杂听觉刺激如何驱动情绪唤醒动态仍不明确。本研究提出一种计算框架,建模自然听觉输入在三个数据集(SEED、LIRIS、自采BAVE)中向动态行为/神经反应的编码过程。基于平行听觉层级的神经生物学原理,将音频分解为多层次特征(通过经典算法及wav2vec 2.0/Hubert),从原始音频与分离的人声/背景音轨中提取,并通过跨数据集分析映射至情绪相关反应。结果表明,来自wav2vec 2.0/Hubert最后一层的高阶语义表征在情绪编码中占主导地位,其与行为标注及多数脑区动态神经同步的映射强度显著高于低阶声学特征(p < 0.05)。值得注意的是,wav2vec 2.0/Hubert中间层(平衡声学-语义信息)在各数据集中表现优于最终层。此外,人声与背景音在不同数据集呈现依赖于刺激能量分布的情绪诱发偏倚(如LIRIS因背景音能量更高而偏好背景音),神经分析显示人声主导前额叶/颞叶活动,背景音则在边缘系统更突出。该研究融合情感计算与神经科学,揭示听觉-情绪编码的分层机制,为自适应情绪感知系统及跨学科音频-情感交互研究提供基础。
原文摘要 · Abstract (English)
In affective neuroscience and emotion-aware AI, understanding how complex auditory stimuli drive emotion arousal dynamics remains unresolved. This study introduces a computational framework to model the brain's encoding of naturalistic auditory inputs into dynamic behavioral/neural responses across three datasets (SEED, LIRIS, self-collected BAVE). Guided by neurobiological principles of parallel auditory hierarchy, we decompose audio into multilevel auditory features (through classical algorithms and wav2vec 2.0/Hubert) from the original and isolated human voice/background soundtrack elements, mapping them to emotion-related responses via cross-dataset analyses. Our analysis reveals that high-level semantic representations (derived from the final layer of wav2vec 2.0/Hubert) exert a dominant role in emotion encoding, outperforming low-level acoustic features with significantly stronger mappings to behavioral annotations and dynamic neural synchrony across most brain regions ($p < 0.05$). Notably, middle layers of wav2vec 2.0/hubert (balancing acoustic-semantic information) surpass the final layers in emotion induction across datasets. Moreover, human voices and soundtracks show dataset-dependent emotion-evoking biases aligned with stimulus energy distribution (e.g., LIRIS favors soundtracks due to higher background energy), with neural analyses indicating voices dominate prefrontal/temporal activity while soundtracks excel in limbic regions. By integrating affective computing and neuroscience, this work uncovers hierarchical mechanisms of auditory-emotion encoding, providing a foundation for adaptive emotion-aware systems and cross-disciplinary explorations of audio-affective interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。