arXiv:2412.19685cs.CVcs.AI2024-12ACL被引 7

为伪造人脸图像生成'哪里被改+为什么这么改'的解释报告

Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline

  • 提出新任务:同时定位篡改区域并生成自然语言解释
  • 构建15万样本数据集,含精准掩码与人工撰写的文本说明
  • 设计统一框架,实现视觉与语言跨模态协同推理

现有面部伪造检测方法多聚焦于二分类或像素级定位,难以提供操作层面的语义解释。为此,本文提出伪造归属报告生成这一新型多模态任务,旨在联合定位伪造区域(何处)并生成基于编辑过程的自然语言解释(为何)。为推动该方向研究,我们构建了大规模多模态篡改追踪数据集MMTT,包含152,217个样本,每个样本配有基于编辑流程的真值掩码和人工撰写的文本描述,确保标注精度与语言丰富性。同时提出ForgeryTalker框架,通过共享编码器(图像编码器+Q-former)与双解码器,实现视觉与语言的端到端联合建模,支持一致的跨模态推理。实验表明,ForgeryTalker在报告生成与伪造定位两个子任务上分别达到59.3 CIDEr和73.67 IoU,为可解释多媒体取证建立基准。数据集与代码将公开以促进后续研究。

原文摘要 · Abstract (English)

Existing facial forgery detection methods typically focus on binary classification or pixel-level localization, providing little semantic insight into the nature of the manipulation. To address this, we introduce Forgery Attribution Report Generation, a new multimodal task that jointly localizes forged regions ("Where") and generates natural language explanations grounded in the editing process ("Why"). This dual-focus approach goes beyond traditional forensics, providing a comprehensive understanding of the manipulation. To enable research in this domain, we present Multi-Modal Tamper Tracing (MMTT), a large-scale dataset of 152,217 samples, each with a process-derived ground-truth mask and a human-authored textual description, ensuring high annotation precision and linguistic richness. We further propose ForgeryTalker, a unified end-to-end framework that integrates vision and language via a shared encoder (image encoder + Q-former) and dual decoders for mask and text generation, enabling coherent cross-modal reasoning. Experiments show that ForgeryTalker achieves competitive performance on both report generation and forgery localization subtasks, i.e., 59.3 CIDEr and 73.67 IoU, respectively, establishing a baseline for explainable multimedia forensics. Dataset and code will be released to foster future research.

伪造检测多模态可解释性图文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。