通过形态学融合局部与全局信息,提升伪造区域定位精度。
Morphology-optimized Multi-Scale Fusion: Combining Local Artifacts and Mesoscopic Semantics for Deepfake Detection and Localization
- 分别从局部细节和全局语义出发独立预测伪造区域
- 采用形态学操作融合结果,抑制噪声并增强空间一致性
- 适合需要精准定位伪造内容的场景,如媒体审核
尽管深度伪造检测的分类准确率不断提升,但精确识别伪造区域仍是难题。现有方法通常在训练中加入伪造区域标注,但常忽视局部细节与全局语义的互补性,且简单拼接局部与全局预测会放大噪声。为此,我们提出一种新方法:分别从局部和全局视角独立预测伪造区域,并使用形态学操作进行融合,有效抑制噪声、增强空间连贯性。大量实验验证了各模块对定位准确性和鲁棒性的提升效果。
原文摘要 · Abstract (English)
While the pursuit of higher accuracy in deepfake detection remains a central goal, there is an increasing demand for precise localization of manipulated regions. Despite the remarkable progress made in classification-based detection, accurately localizing forged areas remains a significant challenge. A common strategy is to incorporate forged region annotations during model training alongside manipulated images. However, such approaches often neglect the complementary nature of local detail and global semantic context, resulting in suboptimal localization performance. Moreover, an often-overlooked aspect is the fusion strategy between local and global predictions. Naively combining the outputs from both branches can amplify noise and errors, thereby undermining the effectiveness of the localization. To address these issues, we propose a novel approach that independently predicts manipulated regions using both local and global perspectives. We employ morphological operations to fuse the outputs, effectively suppressing noise while enhancing spatial coherence. Extensive experiments reveal the effectiveness of each module in improving the accuracy and robustness of forgery localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。