为伪造人脸图像生成'哪里被改+为什么这么改'的解释报告
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
- 提出新任务:同时定位篡改区域并生成自然语言解释
- 构建15万样本数据集,含精准掩码与人工撰写的文本说明
- 设计统一框架,实现视觉与语言跨模态协同推理
现有面部伪造检测方法多聚焦于二分类或像素级定位,难以提供操作层面的语义解释。为此,本文提出伪造归属报告生成这一新型多模态任务,旨在联合定位伪造区域(何处)并生成基于编辑过程的自然语言解释(为何)。为推动该方向研究,我们构建了大规模多模态篡改追踪数据集MMTT,包含152,217个样本,每个样本配有基于编辑流程的真值掩码和人工撰写的文本描述,确保标注精度与语言丰富性。同时提出ForgeryTalker框架,通过共享编码器(图像编码器+Q-former)与双解码器,实现视觉与语言的端到端联合建模,支持一致的跨模态推理。实验表明,ForgeryTalker在报告生成与伪造定位两个子任务上分别达到59.3 CIDEr和73.67 IoU,为可解释多媒体取证建立基准。数据集与代码将公开以促进后续研究。
原文摘要 · Abstract (English)
Existing facial forgery detection methods typically focus on binary classification or pixel-level localization, providing little semantic insight into the nature of the manipulation. To address this, we introduce Forgery Attribution Report Generation, a new multimodal task that jointly localizes forged regions ("Where") and generates natural language explanations grounded in the editing process ("Why"). This dual-focus approach goes beyond traditional forensics, providing a comprehensive understanding of the manipulation. To enable research in this domain, we present Multi-Modal Tamper Tracing (MMTT), a large-scale dataset of 152,217 samples, each with a process-derived ground-truth mask and a human-authored textual description, ensuring high annotation precision and linguistic richness. We further propose ForgeryTalker, a unified end-to-end framework that integrates vision and language via a shared encoder (image encoder + Q-former) and dual decoders for mask and text generation, enabling coherent cross-modal reasoning. Experiments show that ForgeryTalker achieves competitive performance on both report generation and forgery localization subtasks, i.e., 59.3 CIDEr and 73.67 IoU, respectively, establishing a baseline for explainable multimedia forensics. Dataset and code will be released to foster future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。