用语义推理+专业取证分析,精准定位图像篡改区域
Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
- 先用语义理解提出可疑区域,再通过多尺度取证特征验证修正
- 在多个数据集上实现领先精度,对复杂篡改仍保持高鲁棒性
- 适合需要高精度图像伪造检测的安防与媒体审核场景
日益复杂的图像篡改技术亟需可靠的取证方案,以实现精确的篡改定位。现有多模态大模型虽具备上下文理解能力,但难以捕捉细微的低层取证特征。本文提出一种「提出-修正」框架,首先利用适配取证任务的LLaVA模型基于语义理解生成初步分析和定位;随后引入取证修正模块,通过多尺度特征分析与多个专用滤波器集成的技术证据,系统验证并优化初始结果。同时设计增强分割模块,将关键取证线索融入SAM的图像嵌入中,克服其固有的语义偏差,实现精准边界划分。该框架融合先进多模态推理与经典取证方法,确保语义推测经由实证验证与强化,在多个数据集上达到顶尖性能,展现卓越泛化能力与鲁棒性。
原文摘要 · Abstract (English)
The increasing sophistication of image manipulation techniques demands robust forensic solutions that can both reliably detect alterations and precisely localize tampered regions. Recent Multimodal Large Language Models (MLLMs) show promise by leveraging world knowledge and semantic understanding for context-aware detection, yet they struggle with perceiving subtle, low-level forensic artifacts crucial for accurate manipulation localization. This paper presents a novel Propose-Rectify framework that effectively bridges semantic reasoning with forensic-specific analysis. In the proposal stage, our approach utilizes a forensic-adapted LLaVA model to generate initial manipulation analysis and preliminary localization of suspicious regions based on semantic understanding and contextual reasoning. In the rectification stage, we introduce a Forensics Rectification Module that systematically validates and refines these initial proposals through multi-scale forensic feature analysis, integrating technical evidence from several specialized filters. Additionally, we present an Enhanced Segmentation Module that incorporates critical forensic cues into SAM's encoded image embeddings, thereby overcoming inherent semantic biases to achieve precise delineation of manipulated regions. By synergistically combining advanced multimodal reasoning with established forensic methodologies, our framework ensures that initial semantic proposals are systematically validated and enhanced through concrete technical evidence, resulting in comprehensive detection accuracy and localization precision. Extensive experimental validation demonstrates state-of-the-art performance across diverse datasets with exceptional robustness and generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。