用历史证据和领域知识增强视频真假判断,提升跨域检测效果
RASR: Retrieval-Augmented Semantic Reasoning for Fake News Video Detection
- 从动态记忆库中检索相关历史证据,构建跨视频语义关联
- 结合领域先验知识,让大模型生成深度分析报告,准确率提升0.93%
- 适合需要高可信度、跨领域视频真伪识别的场景
多模态虚假视频检测对维护网络信息可信度至关重要。现有方法主要通过构建多模态特征融合表示或利用预训练语言模型分析视频-文本一致性来验证内容真实性,但仍存在两大局限:(1) 缺乏跨实例全局语义关联,难以有效利用历史关联证据验证当前视频;(2) 不同领域间语义差异阻碍通用知识迁移,缺乏领域专家知识引导。为此,我们提出一种新型检索增强语义推理框架RASR。首先,跨实例语义解析与检索器(CSPR)将视频分解为高层语义单元,并从动态记忆库中检索相关关联证据。其次,领域引导多模态推理模块(DGMP)引入领域先验,驱动专家级多模态大语言模型生成领域感知的深度分析报告。最后,多视角特征解耦与融合模块(MVDFF)通过自适应门控机制整合多维特征,实现鲁棒的真实性判定。在FakeSV和FakeTT数据集上的大量实验表明,RASR显著优于现有基线方法,具备更强的跨域泛化能力,整体检测准确率最高提升0.93%。
原文摘要 · Abstract (English)
Multimodal fake news video detection is a crucial research direction for maintaining the credibility of online information. Existing studies primarily verify content authenticity by constructing multimodal feature fusion representations or utilizing pre-trained language models to analyze video-text consistency. However, these methods still face the following limitations: (1) lacking cross-instance global semantic correlations, making it difficult to effectively utilize historical associative evidence to verify the current video; (2) semantic discrepancies across domains hinder the transfer of general knowledge, lacking the guidance of domain-specific expert knowledge. To this end, we propose a novel Retrieval-Augmented Semantic Reasoning (RASR) framework. First, a Cross-instance Semantic Parser and Retriever (CSPR) deconstructs the video into high-level semantic primitives and retrieves relevant associative evidence from a dynamic memory bank. Subsequently, a Domain-Guided Multimodal Reasoning (DGMP) module incorporates domain priors to drive an expert multimodal large language model in generating domain-aware, in-depth analysis reports. Finally, a Multi-View Feature Decoupling and Fusion (MVDFF) module integrates multi-dimensional features through an adaptive gating mechanism to achieve robust authenticity determination. Extensive experiments on the FakeSV and FakeTT datasets demonstrate that RASR significantly outperforms state-of-the-art baselines, achieves superior cross-domain generalization, and improves the overall detection accuracy by up to 0.93%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。