无需训练,用大模型链式推理实现图像篡改检测与定位
Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
- 利用多模态大模型构建可解释的推理链,模拟专家办案流程
- 在多个基准上超越现有无训练方法,媲美有监督模型表现
- 适合需要快速部署且重视结果可解释性的安全检测场景
图像篡改技术的发展带来严重安全威胁,亟需有效的图像篡改定位(IML)方法。尽管有监督的IML性能优异,但依赖昂贵的像素级标注。现有弱监督或无训练方法常表现不佳且缺乏可解释性。本文提出无训练框架In-Context Forensic Chain(ICFC),利用多模态大语言模型(MLLMs)实现可解释的IML任务。ICFC结合对象化规则构建与自适应过滤,建立可靠知识库,并采用多步渐进式推理流程,模拟从粗略提案到精细取证的专家工作流。该设计使MLLM推理能力被系统性用于图像分类、像素级定位及文本级解释。在多个基准测试中,ICFC不仅超越现有最先进无训练方法,还达到甚至超过弱监督和全监督方法的性能。
原文摘要 · Abstract (English)
Advances in image tampering pose serious security threats, underscoring the need for effective image manipulation localization (IML). While supervised IML achieves strong performance, it depends on costly pixel-level annotations. Existing weakly supervised or training-free alternatives often underperform and lack interpretability. We propose the In-Context Forensic Chain (ICFC), a training-free framework that leverages multi-modal large language models (MLLMs) for interpretable IML tasks. ICFC integrates an objectified rule construction with adaptive filtering to build a reliable knowledge base and a multi-step progressive reasoning pipeline that mirrors expert forensic workflows from coarse proposals to fine-grained forensics results. This design enables systematic exploitation of MLLM reasoning for image-level classification, pixel-level localization, and text-level interpretability. Across multiple benchmarks, ICFC not only surpasses state-of-the-art training-free methods but also achieves competitive or superior performance compared to weakly and fully supervised approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。