arXiv:2503.16454cs.HCcs.AI2025-03

基于脑结构对齐的音视频情感生成模型,提升可解释性与效率

An Audio-Visual Fusion Emotion Generation Model Based on Neuroanatomical Alignment

  • 通过模块化设计模拟脑内视听情感通路融合
  • 音视频联合生成的情感相似度显著优于单模态模型
  • 适合追求可解释性的情感计算研究者使用

在情感计算领域,传统情感生成方法主要依赖深度学习和大规模情感数据集。然而,深度学习模型往往复杂且难以解释,而构建标准化的大规模情感数据集成本高昂。为此,我们提出一种新框架——基于神经解剖对齐的音视频融合脑启发情感学习(AVF-BEL)。该框架通过模块化组件集成,改进音视频情感融合与生成模型,实现更轻量、可解释的情感学习过程。其模拟大脑中视觉、听觉与情感通路的整合机制,优化视听模态间的情感特征融合,改进传统脑情绪学习(BEL)模型。实验表明,该模型在音视频联合情感生成上相比单模态视觉或听觉模型具有显著更高的相似度,符合视听刺激协同增强情感生成的基本现象。本工作不仅提升了情感智能的可解释性与效率,也为情感计算技术发展提供了新思路。源代码已开源:https://github.com/OpenHUTB/emotion

原文摘要 · Abstract (English)

In the field of affective computing, traditional methods for generating emotions predominantly rely on deep learning techniques and large-scale emotion datasets. However, deep learning techniques are often complex and difficult to interpret, and standardizing large-scale emotional datasets are difficult and costly to establish. To tackle these challenges, we introduce a novel framework named Audio-Visual Fusion for Brain-like Emotion Learning(AVF-BEL). In contrast to conventional brain-inspired emotion learning methods, this approach improves the audio-visual emotion fusion and generation model through the integration of modular components, thereby enabling more lightweight and interpretable emotion learning and generation processes. The framework simulates the integration of the visual, auditory, and emotional pathways of the brain, optimizes the fusion of emotional features across visual and auditory modalities, and improves upon the traditional Brain Emotional Learning (BEL) model. The experimental results indicate a significant improvement in the similarity of the audio-visual fusion emotion learning generation model compared to single-modality visual and auditory emotion learning and generation model. Ultimately, this aligns with the fundamental phenomenon of heightened emotion generation facilitated by the integrated impact of visual and auditory stimuli. This contribution not only enhances the interpretability and efficiency of affective intelligence but also provides new insights and pathways for advancing affective computing technology. Our source code can be accessed here: https://github.com/OpenHUTB/emotion}{https://github.com/OpenHUTB/emotion.

情感计算音视频融合脑启发模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。