提升图文语义对齐,精准定位多模态篡改内容
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
- 用现成大模型生成伪造图文对,强化跨模态对齐
- 在DGM4数据集上检测准确率显著优于现有方法
- 适合研究多媒体真实性检测与视觉语言模型的学者
我们提出ASAP框架,用于多模态媒体篡改的检测与定位(DGM4)。通过深入分析发现,图像与文本间细粒度的跨模态语义对齐对精准检测篡改至关重要。现有方法普遍忽视这一环节,制约了检测性能的提升。为此,本工作聚焦于增强语义对齐学习以推动该任务发展。具体地,利用现成的多模态大语言模型(MLLMs)和大语言模型(LLMs)构建配对的图像-文本样本,尤其针对被篡改实例;随后进行跨模态对齐学习以提升语义一致性。除显式辅助线索外,还设计了篡改引导交叉注意力(MGCA),为模型提供隐式指导以增强篡改感知能力。训练时借助标注的定位真值,MGCA促使模型更关注篡改区域,抑制正常区域响应,从而提升对篡改内容的捕捉能力。在DGM4数据集上的大量实验表明,所提模型显著超越对比方法。
原文摘要 · Abstract (English)
We present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM4).Upon thorough examination, we observe that accurate fine-grained cross-modal semantic alignment between the image and text is vital for accurately manipulation detection and grounding. While existing DGM4 methods pay rare attention to the cross-modal alignment, hampering the accuracy of manipulation detecting to step further. To remedy this issue, this work targets to advance the semantic alignment learning to promote this task. Particularly, we utilize the off-the-shelf Multimodal Large-Language Models (MLLMs) and Large Language Models (LLMs) to construct paired image-text pairs, especially for the manipulated instances. Subsequently, a cross-modal alignment learning is performed to enhance the semantic alignment. Besides the explicit auxiliary clues, we further design a Manipulation-Guided Cross Attention (MGCA) to provide implicit guidance for augmenting the manipulation perceiving. With the grounding truth available during training, MGCA encourages the model to concentrate more on manipulated components while downplaying normal ones, enhancing the model's ability to capture manipulations. Extensive experiments are conducted on the DGM4 dataset, the results demonstrate that our model can surpass the comparison method with a clear margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。