arXiv:2602.02175cs.CV2026-02被引 1

用粗粒度标注实现图文篡改定位,效果接近全监督方法。

CIEC: Coupling Implicit and Explicit Cues for Multimodal Weakly Supervised Manipulation Localization

  • 融合视觉与文本隐含/显式线索,通过双分支弱监督定位篡改区域。
  • 在多个数据集上达到接近全监督的精度,关键指标提升显著。
  • 适合缺乏细粒度标注资源但需高精度篡改检测的场景。

为应对虚假信息威胁,多模态篡改定位受到广泛关注。现有方法依赖昂贵且耗时的细粒度标注(如像素/词级标注)。本文提出耦合隐式与显式线索的CIEC框架,仅使用图像/句子级粗粒度标注,实现图像-文本对的弱监督篡改定位。包含基于图像和文本的双分支弱监督定位模块。图像分支设计了文本引导的补丁选择(TRPS)模块,结合视觉与文本伪造线索,并利用空间先验锁定可疑区域,再通过背景静音与空间对比约束抑制无关区域干扰。文本分支设计了视觉偏差校准的词定位(VCTG)模块,聚焦语义关键词,利用相对视觉偏置辅助词定位,再通过非对称稀疏性与语义一致性约束缓解标签噪声并保证可靠性。大量实验表明,CIEC在多个评估指标上表现优异,性能接近全监督方法。

原文摘要 · Abstract (English)

To mitigate the threat of misinformation, multimodal manipulation localization has garnered growing attention. Consider that current methods rely on costly and time-consuming fine-grained annotations, such as patch/token-level annotations. This paper proposes a novel framework named Coupling Implicit and Explicit Cues (CIEC), which aims to achieve multimodal weakly-supervised manipulation localization for image-text pairs utilizing only coarse-grained image/sentence-level annotations. It comprises two branches, image-based and text-based weakly-supervised localization. For the former, we devise the Textual-guidance Refine Patch Selection (TRPS) module. It integrates forgery cues from both visual and textual perspectives to lock onto suspicious regions aided by spatial priors. Followed by the background silencing and spatial contrast constraints to suppress interference from irrelevant areas. For the latter, we devise the Visual-deviation Calibrated Token Grounding (VCTG) module. It focuses on meaningful content words and leverages relative visual bias to assist token localization. Followed by the asymmetric sparse and semantic consistency constraints to mitigate label noise and ensure reliability. Extensive experiments demonstrate the effectiveness of our CIEC, yielding results comparable to fully supervised methods on several evaluation metrics.

篡改检测弱监督图文对多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。