arXiv:2604.06728cs.CVcs.AI2026-04中稿 · ICIC 2026

通过建模模态不确定性,提升社交媒体讽刺检测的鲁棒性。

URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection

论文配图:URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
图 1 · 摘自论文原文
  • 用高斯后验建模文本、图像及交互的不确定性
  • 动态调整模态权重,抑制噪声信息干扰
  • 适合处理真实社交数据中模态质量不一的问题

多模态讽刺检测(MSD)旨在通过文本与图像之间的语义不一致识别讽刺意图。尽管现有方法通过跨模态交互和不一致推理提升了性能,但大多将不同模态视为同等可靠。在真实社交媒体帖子中,文本与图像常存在噪声水平和相关性差异,导致确定性融合易受噪声影响并弱化不一致线索。为此,本文提出不确定性感知的鲁棒多模态融合(URMF)框架。URMF首先通过多头交叉注意力将视觉信息注入文本表征,再在融合语义空间中使用自注意力增强不一致推理。它将文本、视觉及交互表征建模为可学习的高斯后验,以估计各模态特异性不确定性,并基于此动态调整融合中的模态贡献,抑制不可靠证据。进一步通过统一目标优化模型,结合信息瓶颈正则化、模态先验正则化、跨模态分布对齐和不确定性驱动的对比学习。在公开的MSD与MMSD2基准上的实验表明,URMF优于代表性单模态、多模态及大语言模型基线。结果证明,显式建模不确定性可同时提升多模态讽刺检测的准确率与鲁棒性。

原文摘要 · Abstract (English)

Multimodal sarcasm detection (MSD) aims to identify sarcastic intent from semantic incongruity between text and image. Although recent methods have improved MSD through cross-modal interaction and incongruity reasoning, most still treat modalities as equally reliable. In real social media posts, however, text and images often differ in noise level and relevance, making deterministic fusion susceptible to noisy evidence and weakened incongruity cues. To address this issue, we propose Uncertainty-aware Robust Multimodal Fusion (URMF), a unified framework for robust MSD. URMF first injects visual evidence into textual representations through multi-head cross-attention, and then applies self-attention in the fused semantic space to enhance incongruity reasoning. It models textual, visual, and interaction-aware representations as learnable Gaussian posteriors to estimate modality-specific uncertainty. Based on the estimated uncertainty, URMF dynamically adjusts modality contributions during fusion to suppress unreliable evidence. We further optimize the model with a unified objective that combines information bottleneck regularization, modality prior regularization, cross-modal distribution alignment, and uncertainty-driven contrastive learning. Experiments on the public MSD and MMSD2 benchmarks show that URMF outperforms representative unimodal, multimodal, and MLLM-based baselines. The results demonstrate that explicit uncertainty modeling can improve both accuracy and robustness in multimodal sarcasm detection.

多模态融合讽刺检测不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。