arXiv:2509.21774cs.CVcs.CY2025-09

无需训练,用图推理提升多模态伪造检测能力

Training-Free Multimodal Deepfake Detection via Graph Reasoning

  • 构建图结构模型捕捉样本间关联,动态筛选关键示例
  • 在四种伪造类型上超越现有方法,无需微调大模型
  • 适合快速部署的伪造检测场景,尤其适合资源受限环境

多模态深度伪造检测(MDD)旨在识别视觉、文本和音频模态中的篡改内容,以增强现代信息系统的可靠性。尽管大型视觉语言模型(LVLM)具备强大的多模态推理能力,但在捕捉细微伪造线索、解决跨模态不一致以及执行任务对齐检索方面仍存在局限。为此,我们提出无需训练的GASP-ICL框架,通过保留语义相关性并注入任务感知知识来增强LVLM。该框架利用适配于MDD的特征提取器检索对齐的图文对,构建候选集;进一步设计图结构泰勒自适应评分器(GSTAS),捕获跨样本关系并传播与查询对齐的信号,生成具有判别性的示范样本。这使得能精准选择语义对齐且任务相关的演示实例,显著提升LVLM在鲁棒性多模态伪造检测中的表现。在四种伪造类型上的实验表明,GASP-ICL在不进行LVLM微调的情况下超越强基线模型。

原文摘要 · Abstract (English)

Multimodal deepfake detection (MDD) aims to uncover manipulations across visual, textual, and auditory modalities, thereby reinforcing the reliability of modern information systems. Although large vision-language models (LVLMs) exhibit strong multimodal reasoning, their effectiveness in MDD is limited by challenges in capturing subtle forgery cues, resolving cross-modal inconsistencies, and performing task-aligned retrieval. To this end, we propose Guided Adaptive Scorer and Propagation In-Context Learning (GASP-ICL), a training-free framework for MDD. GASP-ICL employs a pipeline to preserve semantic relevance while injecting task-aware knowledge into LVLMs. We leverage an MDD-adapted feature extractor to retrieve aligned image-text pairs and build a candidate set. We further design the Graph-Structured Taylor Adaptive Scorer (GSTAS) to capture cross-sample relations and propagate query-aligned signals, producing discriminative exemplars. This enables precise selection of semantically aligned, task-relevant demonstrations, enhancing LVLMs for robust MDD. Experiments on four forgery types show that GASP-ICL surpasses strong baselines, delivering gains without LVLM fine-tuning.

深度伪造检测多模态图神经网络零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。