用法庭辩论思路定位图片篡改区域,提升模糊情况下的判断准确率。
The Courtroom Trial of Pixels: Robust Image Manipulation Localization via Adversarial Evidence and Reinforcement Learning Judgment

- 构建控辩双流模型,分别输出篡改与真实证据。
- 在模糊区域通过强化学习裁判模型重推理,提升定位精度。
- 适合需要高可信度图像真伪检测的场景。
尽管现有图像篡改定位(IML)方法引入真实性监督,但通常仅作为辅助信号增强对篡改痕迹的敏感性,而非显式建模为对抗性证据。当篡改痕迹微弱或经后期处理与噪声干扰后,这些方法难以明确比较篡改与真实证据,导致模糊区域预测不可靠。为此,我们提出一种法庭式裁决框架,将IML任务视为证据交锋与判断过程。框架包含控方流、辩方流和裁判模型。首先在共享多尺度编码器上构建双假设分割架构:控方流主张存在篡改,辩方流主张内容真实。在边缘先验引导下,通过级联多层级融合、双向分歧抑制与动态辩论精炼,生成篡改与真实区域的证据。进一步设计强化学习裁判模型,对不确定区域进行策略性重推断与精炼,输出篡改区域掩码。裁判模型采用基于优势的奖励机制与软IoU目标函数训练,并通过熵与跨假设一致性校准可靠性。实验表明,本模型在平均性能上优于当前最优的IML方法。
原文摘要 · Abstract (English)
Although some existing image manipulation localization (IML) methods incorporate authenticity-related supervision, this information is typically utilized merely as an auxiliary training signal to enhance the model's sensitivity to manipulation artifacts, rather than being explicitly modeled as localization evidence opposing the manipulated regions. Consequently, when manipulation traces are subtle or degraded by post-processing and noise, these methods struggle to explicitly compare manipulated and authentic evidence, resulting in unreliable predictions in ambiguous areas. To address these issues, we propose a courtroom-style adjudication framework that regards IML task as the confrontation of evidence followed by judgment. The framework comprises a prosecution stream, a defense stream, and a judge model. We first build a dual-hypothesis segmentation architecture on a shared multi-scale encoder, in which the prosecution stream asserts manipulation and the defense stream asserts authenticity. Guided by edge priors, it produces evidence for manipulated and authentic regions through cascaded multi-level fusion, bidirectional disagreement suppression, and dynamic debate refinement. We further develop a reinforcement learning judge model that performs strategic re-inference and refinement on uncertain regions, yielding a manipulated-region mask. The judge model is trained with advantage-based rewards and a soft-IoU objective, and reliability is calibrated via entropy and cross-hypothesis consistency. Experimental results show that our model achieves superior average performance compared with SOTA IML methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。