通过关系上下文学习与多路融合,提升多模态讽刺检测准确率。
RCLMuFN: Relational Context Learning and Multiplex Fusion Network for Multimodal Sarcasm Detection
- 构建关系上下文模块,捕捉文本与图像间的动态交互
- 在两个数据集上达到当前最优性能,超越已有方法
- 适合研究多模态情感分析与讽刺识别的学者
讽刺通常以相反于真实意图的方式表达轻蔑或批评情绪。准确识别讽刺有助于过滤网络不良信息,减少恶意诽谤与谣言传播。然而,自动讽刺检测对机器而言仍极具挑战,因其高度依赖复杂的关系上下文。现有方法多关注构建文本与图像间的图结构关系,却忽视了二者间关键的关系上下文信息,且难以建模上下文随时间演变的动态特性,限制了模型泛化能力。为此,本文提出关系上下文学习与多路融合网络(RCLMuFN)。首先,采用四个特征提取器从原始文本与图像中全面提取特征,挖掘潜在被忽略的表征。其次,利用关系上下文学习模块捕获文本与图像的上下文信息,并通过浅层与深层交互捕捉动态变化。最后,通过多路特征融合模块深入整合来自不同交互上下文的多模态特征,增强模型泛化性。在两个多模态讽刺检测数据集上的大量实验表明,所提方法取得当前最优性能。
原文摘要 · Abstract (English)
Sarcasm typically conveys emotions of contempt or criticism by expressing a meaning that is contrary to the speaker's true intent. Accurate detection of sarcasm aids in identifying and filtering undesirable information on the Internet, thereby reducing malicious defamation and rumor-mongering. Nonetheless, the task of automatic sarcasm detection remains highly challenging for machines, as it critically depends on intricate factors such as relational context. Most existing multimodal sarcasm detection methods focus on introducing graph structures to establish entity relationships between text and images while neglecting to learn the relational context between text and images, which is crucial evidence for understanding the meaning of sarcasm. In addition, the meaning of sarcasm changes with the evolution of different contexts, but existing methods may not be accurate in modeling such dynamic changes, limiting the generalization ability of the models. To address the above issues, we propose a relational context learning and multiplex fusion network (RCLMuFN) for multimodal sarcasm detection. Firstly, we employ four feature extractors to comprehensively extract features from raw text and images, aiming to excavate potential features that may have been previously overlooked. Secondly, we utilize the relational context learning module to learn the contextual information of text and images and capture the dynamic properties through shallow and deep interactions. Finally, we employ a multiplex feature fusion module to enhance the generalization of the model by penetratingly integrating multimodal features derived from various interaction contexts. Extensive experiments on two multimodal sarcasm detection datasets show that our proposed method achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。