用记忆检索+推理循环,让视频造假检测既准又可解释。
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
- 通过记忆库检索真实样本,为检测提供上下文锚点。
- 低置信度时自动调用细粒度分析工具,提升判别力。
- 适用于需要透明解释的视频真伪验证场景。
我们针对需兼顾鲁棒性与可解释性的多模态深度伪造检测问题,提出 FakeHunter 框架,融合记忆引导检索、结构化观察-思考-行动推理循环及自适应取证工具调用机制。基于对比语言-图像预训练(CLIP)模型的视觉表征和对比语言-音频预训练(CLAP)模型的音频表征,从大规模记忆库中检索语义对齐的真实样本,为迭代定位与解释可疑篡改提供上下文锚点。当内部置信度较低时,框架会主动触发细粒度分析(如空间区域放大、梅尔频谱图检查),获取判别性证据,而非依赖模糊的边缘分数。我们还发布了 X-AVFake 基准数据集,包含精细标注的伪造类型、受影响区域或实体、推理类别及解释理由,用于强化上下文关联与解释保真度。大量实验表明,FakeHunter 超越现有强基线;消融实验证实,上下文检索与选择性工具激活对提升鲁棒性与解释精度均不可或缺。
原文摘要 · Abstract (English)
We address multimodal deepfake detection requiring both robustness and interpretability by proposing FakeHunter, a unified framework that combines memory guided retrieval, a structured Observation-Thought-Action reasoning loop, and adaptive forensic tool invocation. Visual representations from a Contrastive Language-Image Pretraining (CLIP) model and audio representations from a Contrastive Language-Audio Pretraining (CLAP) model retrieve semantically aligned authentic exemplars from a large scale memory, providing contextual anchors that guide iterative localization and explanation of suspected manipulations. Under low internal confidence the framework selectively triggers fine grained analyses such as spatial region zoom and mel spectrogram inspection to gather discriminative evidence instead of relying on opaque marginal scores. We also release X-AVFake, a comprehensive audio visual forgery benchmark with fine grained annotations of manipulation type, affected region or entity, reasoning category, and explanatory justification, designed to stress contextual grounding and explanation fidelity. Extensive experiments show that FakeHunter surpasses strong multimodal baselines, and ablation studies confirm that both contextual retrieval and selective tool activation are indispensable for improved robustness and explanatory precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。