用常识知识与跨模态相似性提升图文讽刺识别准确率
SemIRNet: A Semantic Irony Recognition Network for Multimodal Sarcasm Detection
- 引入ConceptNet增强常识推理,理解图文隐含关联
- 在词级和样本级设计双粒度语义相似度模块
- 对比学习优化特征分布,提升正负样本区分度
针对多模态讽刺识别中图形与文本隐含关联难以准确捕捉的问题,本文提出语义讽刺识别网络(SemIRNet)。模型有三方面创新:(1)首次引入ConceptNet知识库获取概念知识,增强模型常识推理能力;(2)设计词级与样本级两个跨模态语义相似度检测模块,建模图文在不同粒度上的关联;(3)引入对比学习损失函数,优化样本特征的空间分布,提升正负样本可分性。在公开的多模态讽刺检测基准数据集上,该模型准确率和F1值分别达到88.87%和86.33%,较现有最优方法提升1.64%和2.88%。消融实验验证了知识融合与语义相似度检测对性能提升的关键作用。
原文摘要 · Abstract (English)
Aiming at the problem of difficulty in accurately identifying graphical implicit correlations in multimodal irony detection tasks, this paper proposes a Semantic Irony Recognition Network (SemIRNet). The model contains three main innovations: (1) The ConceptNet knowledge base is introduced for the first time to acquire conceptual knowledge, which enhances the model's common-sense reasoning ability; (2) Two cross-modal semantic similarity detection modules at the word level and sample level are designed to model graphic-textual correlations at different granularities; and (3) A contrastive learning loss function is introduced to optimize the spatial distribution of the sample features, which improves the separability of positive and negative samples. Experiments on a publicly available multimodal irony detection benchmark dataset show that the accuracy and F1 value of this model are improved by 1.64% and 2.88% to 88.87% and 86.33%, respectively, compared with the existing optimal methods. Further ablation experiments verify the important role of knowledge fusion and semantic similarity detection in improving the model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。