提出因果框架联合解决多模态讽刺识别与解释问题
MuVaC: A Variational Causal Framework for Multimodal Sarcasm Understanding in Dialogues
- 构建变分因果模型,模拟人类理解讽刺的认知过程
- 在公开数据集上同时提升讽刺识别与解释效果
- 适合关注多模态情感分析与可解释AI的研究者
社交媒体中多模态对话的讽刺现象普遍,准确理解其真实意图是重要但具挑战性的任务。全面的讽刺分析需兼顾多模态讽刺检测(MSD)和多模态讽刺解释(MuSE)。直观上,检测结果源于解释推理过程。现有研究多将两者作为独立任务处理,即使有整合尝试,也常忽视其内在因果依赖。为此,我们提出MuVaC——一种模仿人类认知机制的变分因果推断框架,实现对多模态特征的鲁棒学习,从而联合优化MSD与MuSE。具体而言,我们从结构因果模型视角建模二者关系,建立变分因果路径以定义联合优化目标;设计先对齐后融合的方法整合多模态特征,生成鲁棒的融合表示用于检测与解释;最后通过确保检测结果与解释的一致性,增强推理可信度。实验表明,MuVaC在多个公开数据集上表现更优,为多模态讽刺理解提供了新视角。
原文摘要 · Abstract (English)
The prevalence of sarcasm in multimodal dialogues on the social platforms presents a crucial yet challenging task for understanding the true intent behind online content. Comprehensive sarcasm analysis requires two key aspects: Multimodal Sarcasm Detection (MSD) and Multimodal Sarcasm Explanation (MuSE). Intuitively, the act of detection is the result of the reasoning process that explains the sarcasm. Current research predominantly focuses on addressing either MSD or MuSE as a single task. Even though some recent work has attempted to integrate these tasks, their inherent causal dependency is often overlooked. To bridge this gap, we propose MuVaC, a variational causal inference framework that mimics human cognitive mechanisms for understanding sarcasm, enabling robust multimodal feature learning to jointly optimize MSD and MuSE. Specifically, we first model MSD and MuSE from the perspective of structural causal models, establishing variational causal pathways to define the objectives for joint optimization. Next, we design an alignment-then-fusion approach to integrate multimodal features, providing robust fusion representations for sarcasm detection and explanation generation. Finally, we enhance the reasoning trustworthiness by ensuring consistency between detection results and explanations. Experimental results demonstrate the superiority of MuVaC in public datasets, offering a new perspective for understanding multimodal sarcasm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。