用视觉问答让假图检测可解释,不训练也能准。
TruthLens:A Training-Free Paradigm for DeepFake Detection
- 把假图检测当视觉问答做,用大模型看图说话
- 在多个难数据集上准确率超传统方法
- 适合需要知道为什么是假图的研究和审核人员
先进AI生成的合成图像泛滥,给识别视觉篡改带来挑战。现有假图检测方法多依赖二分类模型,侧重准确率而忽视可解释性,用户无法理解判断依据。为此,我们提出TruthLens,一种无需训练的新框架,将深度伪造检测重构为视觉问答(VQA)任务。该框架利用先进的大视觉语言模型(LVLMs)观察并描述视觉伪影,并结合GPT-4等大语言模型(LLMs)的推理能力,分析并聚合证据以做出判断。通过多模态融合,TruthLens不仅能分类图像真伪,还能提供可解释的决策理由,增强信任并揭示合成内容的特征线索。大量评估表明,TruthLens在复杂数据集上表现优于传统方法,兼具高精度与强可解释性。通过将检测过程变为推理驱动,TruthLens为应对视觉虚假信息树立了新范式。
原文摘要 · Abstract (English)
The proliferation of synthetic images generated by advanced AI models poses significant challenges in identifying and understanding manipulated visual content. Current fake image detection methods predominantly rely on binary classification models that focus on accuracy while often neglecting interpretability, leaving users without clear insights into why an image is deemed real or fake. To bridge this gap, we introduce TruthLens, a novel training-free framework that reimagines deepfake detection as a visual question-answering (VQA) task. TruthLens utilizes state-of-the-art large vision-language models (LVLMs) to observe and describe visual artifacts and combines this with the reasoning capabilities of large language models (LLMs) like GPT-4 to analyze and aggregate evidence into informed decisions. By adopting a multimodal approach, TruthLens seamlessly integrates visual and semantic reasoning to not only classify images as real or fake but also provide interpretable explanations for its decisions. This transparency enhances trust and provides valuable insights into the artifacts that signal synthetic content. Extensive evaluations demonstrate that TruthLens outperforms conventional methods, achieving high accuracy on challenging datasets while maintaining a strong emphasis on explainability. By reframing deepfake detection as a reasoning-driven process, TruthLens establishes a new paradigm in combating synthetic media, combining cutting-edge performance with interpretability to address the growing threats of visual disinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。