arXiv:2603.14992cs.AIcs.MM2026-03

通过挖掘多模态一致性差异,提升短视频假新闻检测精度。

Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos

  • 设计多粒度跨模态一致性建模机制,捕捉图文音间潜在矛盾。
  • 在中英文数据集上实现比非VLM基线更高的准确率,接近VLM水平。
  • 推理速度提升18-27倍,显存节省93%,适合实际部署场景。

短时视频平台是新闻传播的重要渠道,但也成为多模态虚假信息的温床——各模态单独看似合理,但跨模态关系存在细微不一致,如画面与字幕不符。在两个基准数据集FakeSV(中文)和FakeTT(英文)上,我们观察到明显不对称:真实视频具有高文本-视觉一致性但中等文本-音频一致性,而虚假视频则相反。此外,单一全局一致性得分可作为可解释轴,平滑反映虚假概率与预测误差变化。受此启发,我们提出MAGIC3(模态对抗门控交互与一致性中心分类器),显式建模并暴露多粒度跨三模态一致性信号。MAGIC3结合成对与全局一致性建模,利用跨模态注意力提取帧级与标记级一致性特征,采用多风格大语言模型重写以获得风格鲁棒的文本表示,并引入不确定性感知分类器实现选择性视觉语言模型路由。使用预提取特征,MAGIC3在FakeSV与FakeTT上持续超越最强非VLM基线;在匹配VLM级准确率的同时,两阶段系统实现18–27倍吞吐量提升与93%显存节省,提供显著的成本-性能权衡。

原文摘要 · Abstract (English)

Short-form video platforms are major channels for news but also fertile ground for multimodal misinformation where each modality appears plausible alone yet cross-modal relationships are subtly inconsistent, like mismatched visuals and captions. On two benchmark datasets, FakeSV (Chinese) and FakeTT (English), we observe a clear asymmetry: real videos exhibit high text-visual but moderate text-audio consistency, while fake videos show the opposite pattern. Moreover, a single global consistency score forms an interpretable axis along which fake probability and prediction errors vary smoothly. Motivated by these observations, we present MAGIC3 (Modal-Adversarial Gated Interaction and Consistency-Centric Classifier), a detector that explicitly models and exposes cross-tri-modal consistency signals at multiple granularities. MAGIC3 combines explicit pairwise and global consistency modeling with token- and frame-level consistency signals derived from cross-modal attention, incorporates multi-style LLM rewrites to obtain style-robust text representations, and employs an uncertainty-aware classifier for selective VLM routing. Using pre-extracted features, MAGIC3 consistently outperforms the strongest non-VLM baselines on FakeSV and FakeTT. While matching VLM-level accuracy, the two-stage system achieves 18-27x higher throughput and 93% VRAM savings, offering a strong cost-performance tradeoff.

假新闻检测多模态分析一致性建模高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。