arXiv:2605.02447cs.CLcs.AI2026-05

通过极性调制注意力建模多模态讽刺中的语义不一致,提升识别准确率。

PC-MNet: Dual-Level Congruity Modeling for Multimodal Sarcasm Detection via Polarity-Modulated Attention

论文配图:PC-MNet: Dual-Level Congruity Modeling for Multimodal Sarcasm Detection via Polarity-Modulated Attention
图 1 · 摘自论文原文
  • 设计双层级一致性建模,用极性调制注意力捕捉文本与非语言信号的矛盾
  • 在MUStARD数据集上比最强基线高出3.14%的宏平均F1值
  • 适合研究讽刺识别、多模态情感分析的学者与开发者

多模态讽刺检测旨在精准识别字面语义与非语言线索之间的语用不一致,近年来备受关注。现有方法多依赖朴素的相似性注意力机制和统一的晚期融合策略。由于功能纠缠限制了传统晚期融合,本文引入标量一致性路由机制和先验引导的上下文图,通过不一致感知对比学习驱动两阶段非对称优化,锚定广义不一致流形,仅选择最具判别性的多粒度证据进行融合。在\texttt{MUStARD}基准及其去伪相关性平衡数据集上的大量实验表明,该方法达到新最优性能,相比最强多模态基线在宏平均F1上提升3.14%。通过解耦原子级、组合级与上下文级冲突,为建模人类交流中的微妙语用不一致提供了稳健、解耦的范式。

原文摘要 · Abstract (English)

Multimodal sarcasm detection, which aims to precisely identify pragmatic incongruities between literal text and nonverbal cues, has gained substantial attention in multimodal understanding. Recent advancements have predominantly relied on na\"ıve similarity-based attention mechanisms and uniform late fusion strategies.Furthermore, given that functional entanglement restricts traditional late fusions, we incorporate a scalar congruity routing mechanism and a prior-guided contextual graph. This mechanism anchors a generalized incongruity manifold through a two-stage asymmetric optimization driven by inconsistency-aware contrastive learning, selectively fusing only the most discriminative multi-granularity evidence. Extensive experiments on the \texttt{MUStARD} benchmark and its spurious-correlation-mitigated balanced datasets demonstrate that our approach achieves new state-of-the-art performance, surpassing the strongest multimodal baseline by a substantial 3.14\% improvement in Macro-F1. By architecturally isolating atomic, composition, and contextual conflicts. This work provides a robust, decoupled paradigm for modeling subtle pragmatic incongruities in human communication.

讽刺检测多模态注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。