arXiv:2409.07388cs.CL2024-09被引 23

从NLP视角梳理多模态情感计算最新进展,系统归纳四类核心任务。

Recent Advances in Multimodal Affective Computing: An NLP Perspective

  • 按文本主导场景统一梳理四类任务的定义与方法范式。
  • 对比主流数据集与评估标准,揭示不同任务间共性与差异。
  • 适合关注多模态情感分析、跨模态理解的研究者参考。

多模态情感计算因在理解人类行为与意图中的广泛应用而备受关注,尤其在以文本为核心的多模态场景中。现有研究涵盖多种任务、模态与建模范式,但缺乏统一视角。本文从自然语言处理(NLP)视角系统综述近期进展,聚焦四大代表性任务:多模态情感分析(MSA)、对话中多模态情绪识别(MERC)、多模态基于观点的情感分析(MABSA)以及多模态多标签情绪识别(MMER)。通过比较任务定义、基准数据集与评估协议,将代表性方法归纳为多任务学习、预训练模型、知识增强和上下文建模等关键范式。进一步拓展至面部、语音、生理模态及情绪原因分析等方向。最后指出关键挑战并展望未来方向。为促进后续研究,我们发布了一个精选论文与资源库。

原文摘要 · Abstract (English)

Multimodal affective computing has gained increasing attention due to its broad applications in understanding human behavior and intentions, particularly in text-centric multimodal scenarios. Existing research spans diverse tasks, modalities, and modeling paradigms, yet lacks a unified perspective. In this survey, we systematically review recent advances from an NLP perspective, focusing on four representative tasks: multimodal sentiment analysis (MSA), multimodal emotion recognition in conversation (MERC), multimodal aspect-based sentiment analysis (MABSA), and multimodal multi-label emotion recognition (MMER). We present a unified view by comparing task formulations, benchmark datasets, and evaluation protocols, and by organizing representative methods into key paradigms, including multitask learning, pre-trained models, knowledge enhancement, and contextual modeling. We further extend the discussion to related directions, such as facial, acoustic, and physiological modalities, as well as emotion cause analysis. Finally, we highlight key challenges and outline promising future directions. To facilitate further research, we release a curated repository of relevant works and resources \footnote{https://anonymous.4open.science/r/Multimodal-Affective-Computing-Survey-9819}.

情感计算多模态NLP综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。