融合外部证据与伪造特征,提升多模态谣言检测精度
Multimodal rumor detection enhanced by external evidence and forgery features
- 引入傅里叶变换提取图像频域痕迹与压缩伪影作为伪造特征
- 在微博和推特数据集上F1值显著优于主流基线模型
- 使用BLIP生成精准图像描述,构建跨模态语义对齐桥梁
社交媒体中图文混合内容传播日益普遍,但谣言常利用细微不一致与伪造内容,仅依赖帖子内容难以检测。深度语义错配谣言表面图文一致,威胁网络舆论安全。现有方法虽增强跨模态建模,但受限于特征提取不足、对齐噪声大、融合策略僵化,且忽略验证复杂谣言所需的外部事实证据。为此,我们提出一种融合外部证据与伪造特征的多模态谣言检测模型。该模型采用ResNet34视觉编码器、BERT文本编码器及伪造特征模块,通过傅里叶变换提取频域痕迹与压缩伪影。针对生成式视觉语言模型生成冗长且风格不一的描述导致语义噪声的问题,我们采用专门预训练于视觉语言对齐的BLIP,生成简洁、忠实于图像、风格接近新闻文本的描述,作为可靠的跨模态语义桥梁。设计基于对比学习的BLIP驱动语义对齐模块,联合优化图文与文本描述的对比损失,捕捉视觉与语义层面的不一致。引入门控自适应特征缩放融合机制,动态调节多模态融合,减少冗余。在微博与推特数据集上的实验表明,本模型在准确率、召回率与F1分数上均优于主流基线。
原文摘要 · Abstract (English)
Social media increasingly disseminates information through mixed image text posts, but rumors often exploit subtle inconsistencies and forged content, making detection based solely on post content difficult. Deep semantic mismatch rumors, which superficially align images and texts, pose particular challenges and threaten online public opinion. Existing multimodal rumor detection methods improve cross modal modeling but suffer from limited feature extraction, noisy alignment, and inflexible fusion strategies, while ignoring external factual evidence necessary for verifying complex rumors. To address these limitations, we propose a multimodal rumor detection model enhanced with external evidence and forgery features. The model uses a ResNet34 visual encoder, a BERT text encoder, and a forgery feature module extracting frequency domain traces and compression artifacts via Fourier transformation. While some existing approaches employ large scale generative vision language models for caption generation, their open ended generation tendency produces verbose and stylistically inconsistent descriptions that introduce semantic noise and risk drifting from actual image content. To overcome this, we adopt BLIP specifically pre trained for vision language alignment which generates concise, image faithful descriptions stylistically closer to news text, serving as a reliable semantic bridge across modalities. A BLIP Driven Semantic Alignment Module jointly optimizes text image and text description contrastive losses, capturing inconsistencies at both visual and semantic levels. A gated adaptive feature scaling fusion mechanism dynamically adjusts multimodal fusion and reduces redundancy. Experiments on Weibo and Twitter datasets demonstrate that our model outperforms mainstream baselines in,recall,and F1 score
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。