通过多模态融合预测视频引发的愉悦感,提升模型可解释性。
Modeling Induced Pleasure through Cognitive Appraisal Prediction via Multimodal Fusion

- 结合认知评价理论与模糊模型,用Transformer捕捉多模态动态特征。
- 在愉悦度预测上达到0.6624的峰值准确率,突破语义鸿沟。
- 适合情感计算、智能内容推荐和数字媒体影响研究者使用。
多模态情感计算旨在分析用户生成的社交媒体内容以预测情绪状态。然而,视觉内容如何影响认知解读并引发特定情感体验(如愉悦)仍存在关键空白。本文提出一种新计算模型,通过认知评价变量推断视频诱发的愉悦感。该模型解决四大挑战:(1) 人类标注噪声大且不一致;(2) “正向情绪”与“愉悦”之间的语义差异;(3) 缺乏专门针对愉悦的数据集;(4) 现有黑箱融合方法可解释性差。方法融合数据驱动与认知理论驱动范式,引入认知评价理论与模糊模型构建创新框架。采用基于Transformer的架构与注意力机制实现细粒度多模态特征提取及可解释融合,捕捉愉悦相关的模态间与模态内动态。由此可预测潜在评价变量,弥合语义差距,超越传统统计关联,增强模型可解释性。实验验证方法有效性,在视频诱发愉悦度预测中达到0.6624的峰值准确率。研究为情感内容推荐、智能媒体创作及理解数字媒体对人类情绪的影响提供新思路。
原文摘要 · Abstract (English)
Multimodal affective computing analyzes user-generated social media content to predict emotional states. However, a critical gap remains in understanding how visual content shapes cognitive interpretations and elicits specific affective experiences such as pleasure. This study introduces a novel computational model to infer video-induced pleasure via cognitive appraisal variables. The proposed model addresses four challenges: (1) noisy and inconsistent human labels, (2) the semantic gap between "positive emotions" and "pleasure," (3) the scarcity of pleasure-specific datasets, and (4) the limited interpretability of existing black-box fusion methods. Our approach integrates data-driven and cognitive theory-driven methods, using cognitive appraisal theory and a fuzzy model within an innovative framework. The model employs transformer-based architectures and attention mechanisms for fine-grained multimodal feature extraction and interpretable fusion to capture both inter- and intra-modal dynamics associated with pleasure. This enables the prediction of underlying appraisal variables, thereby bridging the semantic gap and enhancing model explainability beyond conventional statistical associations. Experimental results validate the efficacy of the proposed method in detecting video-induced pleasure, achieving a peak accuracy of 0.6624 in predicting pleasure levels. These findings highlight promising implications for affective content recommendation, intelligent media creation, and advancing our understanding of how digital media influences human emotions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。