融合宏观与微观信息,用新架构提升图像篡改定位精度
Mesoscopic Insights: Orchestrating Multi-scale & Hybrid Architecture for Image Manipulation Localization
- 并行使用Transformer(抓宏观)和CNN(捕微观),构建多尺度表征
- 在4个数据集上超越现有方法,在准确率与效率上均表现更优
- 适合关注图像真实性验证、多尺度特征融合的研究者
介观层面介于宏观与微观之间,填补了两者间的认知空白。图像篡改定位(IML)作为识别虚假图像的关键技术,长期依赖低层级(微观)痕迹。然而实际篡改常针对图像语义,发生在对象级别(宏观层面),其重要性不亚于微观痕迹。因此,将两者整合至介观层面,为IML研究提供新视角。受此启发,本文提出Mesorch架构,实现微观与宏观信息的协同建模:1)并行结合Transformer(提取宏观信息)与CNN(捕捉微观细节);2)跨尺度评估,无缝融合微宏观信息。基于该架构,设计两个基线模型用于解决IML任务。在四个数据集上的大量实验表明,所提模型在性能、计算复杂度与鲁棒性方面均优于当前最先进方法。
原文摘要 · Abstract (English)
The mesoscopic level serves as a bridge between the macroscopic and microscopic worlds, addressing gaps overlooked by both. Image manipulation localization (IML), a crucial technique to pursue truth from fake images, has long relied on low-level (microscopic-level) traces. However, in practice, most tampering aims to deceive the audience by altering image semantics. As a result, manipulation commonly occurs at the object level (macroscopic level), which is equally important as microscopic traces. Therefore, integrating these two levels into the mesoscopic level presents a new perspective for IML research. Inspired by this, our paper explores how to simultaneously construct mesoscopic representations of micro and macro information for IML and introduces the Mesorch architecture to orchestrate both. Specifically, this architecture i) combines Transformers and CNNs in parallel, with Transformers extracting macro information and CNNs capturing micro details, and ii) explores across different scales, assessing micro and macro information seamlessly. Additionally, based on the Mesorch architecture, the paper introduces two baseline models aimed at solving IML tasks through mesoscopic representation. Extensive experiments across four datasets have demonstrated that our models surpass the current state-of-the-art in terms of performance, computational complexity, and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。