arXiv:2510.20531cs.CVcs.AI2025-10被引 2

让AI精准定位人脸伪造区域并给出有依据的解释。

Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis

  • 构建细粒度人脸区域概念树,提升标注可靠性。
  • 生成带分割掩码的文本解释,实现图文对应。
  • 支持任意面部区域提问,适合需要可解释性的研究者。

多模态大模型的发展促进了视觉与语言任务的融合,推动了可解释性深度伪造分析(XDFA)的进步。然而现有方法在细粒度感知上存在不足:数据标注中的伪影描述不可靠且粗糙,模型无法输出文本解释与视觉证据之间的关联,也无法针对任意面部区域进行查询,导致结果缺乏人脸视觉上下文(Facext)的支撑。为此,我们提出Fake-in-Facext(FiFa)框架,重点在于数据标注与模型构建。首先定义面部图像概念树(FICT),将人脸划分为细粒度区域概念,建立更可靠的标注流程FiFa-Annotator。基于此,引入新的伪影接地解释(AGE)任务,生成包含操作区域分割掩码的文本解释。设计统一的多任务学习架构FiFa-MLLM,同时支持丰富的多模态输入与输出,实现细粒度可解释性深度伪造分析。通过多个辅助监督任务,FiFa-MLLM在AGE任务上优于强基线,并在现有XDFA数据集上达到最先进性能。代码与数据将开源于https://github.com/lxq1000/Fake-in-Facext。

原文摘要 · Abstract (English)

The advancement of Multimodal Large Language Models (MLLMs) has bridged the gap between vision and language tasks, enabling the implementation of Explainable DeepFake Analysis (XDFA). However, current methods suffer from a lack of fine-grained awareness: the description of artifacts in data annotation is unreliable and coarse-grained, and the models fail to support the output of connections between textual forgery explanations and the visual evidence of artifacts, as well as the input of queries for arbitrary facial regions. As a result, their responses are not sufficiently grounded in Face Visual Context (Facext). To address this limitation, we propose the Fake-in-Facext (FiFa) framework, with contributions focusing on data annotation and model construction. We first define a Facial Image Concept Tree (FICT) to divide facial images into fine-grained regional concepts, thereby obtaining a more reliable data annotation pipeline, FiFa-Annotator, for forgery explanation. Based on this dedicated data annotation, we introduce a novel Artifact-Grounding Explanation (AGE) task, which generates textual forgery explanations interleaved with segmentation masks of manipulated artifacts. We propose a unified multi-task learning architecture, FiFa-MLLM, to simultaneously support abundant multimodal inputs and outputs for fine-grained Explainable DeepFake Analysis. With multiple auxiliary supervision tasks, FiFa-MLLM can outperform strong baselines on the AGE task and achieve SOTA performance on existing XDFA datasets. The code and data will be made open-source at https://github.com/lxq1000/Fake-in-Facext.

可解释性深度伪造多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。