TruthLens让AI能解释为何判断一张图是伪造的,还能指出具体哪部分有问题。
TruthLens: Visual Grounding for Universal DeepFake Reasoning
- 用多模态大模型结合视觉特征,实现局部区域精准溯源
- 在多个数据集上检测准确率超越现有方法,对新类型伪造也有效
- 支持细粒度提问,如‘眼睛/鼻子是否自然’,解释更透明
随着AI图像生成技术普及,深度伪造内容日益泛滥,传统检测方法多局限于真假二分类且缺乏可解释性。为此,我们提出TruthLens——一种统一、通用性强的新框架,突破二分类限制,提供详尽的文本化推理过程。不同于常规方法,TruthLens采用基于任务的表示融合策略,将多模态大语言模型(MLLM)提供的全局语义信息与仅视觉模型经跨模态适配后提取的局部取证线索相结合,实现对人脸篡改和全合成内容的精细化区域定位推理,支持如“眼睛/鼻子/嘴巴是否真实”等细粒度查询。在多个数据集上的实验表明,TruthLens在检测精度与可解释性方面均达到新基准,能泛化至已见及未见的伪造类型。通过整合高层场景理解与细粒度区域定位,该框架实现了透明化的深度伪造分析,填补了当前研究的关键空白。
原文摘要 · Abstract (English)
Detecting DeepFakes has become a crucial research area as the widespread use of AI image generators enables the effortless creation of face-manipulated and fully synthetic content, while existing methods are often limited to binary classification (real vs. fake) and lack interpretability. To address these challenges, we propose TruthLens, a novel, unified, and highly generalizable framework that goes beyond traditional binary classification, providing detailed, textual reasoning for its predictions. Distinct from conventional methods, TruthLens performs MLLM grounding. TruthLens uses a task-driven representation integration strategy that unites global semantic context from a multimodal large language model (MLLM) with region-specific forensic cues through explicit cross-modal adaptation of a vision-only model. This enables nuanced, region-grounded reasoning for both face-manipulated and fully synthetic content, and supports fine-grained queries such as "Does the eyes/nose/mouth look real or fake?"- capabilities beyond pretrained MLLMs alone. Extensive experiments across diverse datasets demonstrate that TruthLens sets a new benchmark in both forensic interpretability and detection accuracy, generalizing to seen and unseen manipulations alike. By unifying high-level scene understanding with fine-grained region grounding, TruthLens delivers transparent DeepFake forensics, bridging a critical gap in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。