arXiv:2508.11197cs.CLcs.AI2025-08被引 1

通过事件级建模与时间一致性机制,提升多模态假信息检测效果。

E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection

  • 按文本相似性与时间接近度聚类成伪事件,统一处理跨模态信息
  • 融合自注意力与交叉注意力,生成内容感知的上下文嵌入表示
  • 引入动态加权与时间一致性正则,有效缓解类别不平衡问题

社交媒体上的多模态假信息检测面临模态不一致、时间模式变化和严重类别不平衡等挑战。现有方法多独立处理帖子,未能捕捉跨时间与模态的事件级结构。本文提出 E-CaTCH,一个可解释且可扩展的鲁棒假信息检测框架。当需要时,该框架基于文本相似性和时间接近度将帖子聚类为伪事件,并独立处理每个事件。在事件内,使用预训练 BERT 和 ResNet 编码器提取文本与视觉特征,经由模态内自注意力优化后,通过双向交叉注意力对齐。软门控机制融合表示,形成上下文感知的内容嵌入。为建模时间演化,E-CaTCH 将事件划分为重叠时间窗口,采用增强语义漂移与动量信号的趋势感知 LSTM 编码叙事进展。分类在事件层面进行,更契合真实假信息传播动态。为应对类别不平衡并促进稳定学习,模型集成自适应类别加权、时间一致性正则化与难例挖掘,总损失在所有事件上聚合。在 Fakeddit、IND 与 COVID-19 MISINFOGRAPH 数据集上的实验表明,E-CaTCH 持续优于当前最优基线。跨数据集评估进一步验证其鲁棒性、泛化能力与实际适用性。

原文摘要 · Abstract (English)

Detecting multimodal misinformation on social media remains challenging due to inconsistencies between modalities, changes in temporal patterns, and substantial class imbalance. Many existing methods treat posts independently and fail to capture the event-level structure that connects them across time and modality. We propose E-CaTCH, an interpretable and scalable framework for robustly detecting misinformation. If needed, E-CaTCH clusters posts into pseudo-events based on textual similarity and temporal proximity, then processes each event independently. Within each event, textual and visual features are extracted using pre-trained BERT and ResNet encoders, refined via intra-modal self-attention, and aligned through bidirectional cross-modal attention. A soft gating mechanism fuses these representations to form contextualized, content-aware embeddings of each post. To model temporal evolution, E-CaTCH segments events into overlapping time windows and uses a trend-aware LSTM, enhanced with semantic shift and momentum signals, to encode narrative progression over time. Classification is performed at the event level, enabling better alignment with real-world misinformation dynamics. To address class imbalance and promote stable learning, the model integrates adaptive class weighting, temporal consistency regularization, and hard-example mining. The total loss is aggregated across all events. Extensive experiments on Fakeddit, IND, and COVID-19 MISINFOGRAPH demonstrate that E-CaTCH consistently outperforms state-of-the-art baselines. Cross-dataset evaluations further demonstrate its robustness, generalizability, and practical applicability across diverse misinformation scenarios.

假信息检测多模态分析时间建模类别不平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。