arXiv:2604.10132cs.CVcs.AI2026-04

定位图像中改变语义的细微编辑,提升真实场景下的篡改检测能力。

Semantic Manipulation Localization

  • 基于语义敏感性构建三阶段推理框架,捕捉视觉一致下的细微修改。
  • 在自建细粒度基准上显著优于现有方法,定位结果更完整且语义连贯。
  • 适合关注生成图像真实性、语义级篡改检测的研究者与应用开发者。

图像篡改定位(IML)旨在识别图像中被编辑的区域。然而,随着现代图像编辑和生成模型的广泛应用,许多篡改不再产生明显的低层伪影,而是通过微妙但改变物体属性、状态或关系的语义级修改实现,同时保持与周围内容的高度一致性。这使得依赖伪影检测的传统IML方法效果下降。为此,我们提出语义篡改定位(SML)新任务,专注于定位显著改变图像解释的细微语义编辑。我们进一步构建了一个专用的细粒度基准,采用语义驱动的篡改流程并提供像素级标注。基于此任务,我们提出TRACE(目标化认知属性编辑推理)框架,通过三个逐步耦合的组件建模语义敏感性:语义锚定、语义扰动感知和语义约束推理。具体而言,TRACE首先识别支持图像理解的语义有意义区域,然后注入对扰动敏感的频率线索以捕捉强视觉一致性下的细微编辑,并最终通过语义内容与语义范围的联合推理验证候选区域。大量实验表明,TRACE在我们的基准上持续优于现有IML方法,生成更完整、紧凑且语义连贯的定位结果。这些结果证明了超越基于伪影的定位的必要性,并为复杂语义编辑场景下的图像取证提供了新方向。

原文摘要 · Abstract (English)

Image Manipulation Localization (IML) aims to identify edited regions in an image. However, with the increasing use of modern image editing and generative models, many manipulations no longer exhibit obvious low-level artifacts. Instead, they often involve subtle but meaning-altering edits to an object's attributes, state, or relationships while remaining highly consistent with the surrounding content. This makes conventional IML methods less effective because they mainly rely on artifact detection rather than semantic sensitivity. To address this issue, we introduce Semantic Manipulation Localization (SML), a new task that focuses on localizing subtle semantic edits that significantly change image interpretation. We further construct a dedicated fine-grained benchmark for SML using a semantics-driven manipulation pipeline with pixel-level annotations. Based on this task, we propose TRACE (Targeted Reasoning of Attributed Cognitive Edits), an end-to-end framework that models semantic sensitivity through three progressively coupled components: semantic anchoring, semantic perturbation sensing, and semantic-constrained reasoning. Specifically, TRACE first identifies semantically meaningful regions that support image understanding, then injects perturbation-sensitive frequency cues to capture subtle edits under strong visual consistency, and finally verifies candidate regions through joint reasoning over semantic content and semantic scope. Extensive experiments show that TRACE consistently outperforms existing IML methods on our benchmark and produces more complete, compact, and semantically coherent localization results. These results demonstrate the necessity of moving beyond artifact-based localization and provide a new direction for image forensics in complex semantic editing scenarios.

图像伪造语义分析定位检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。