聚焦人脸局部区域,提升多模态伪造检测精度。
M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
- 基于BLIP-2提取局部特征,融合人脸先验知识
- 设计对齐融合模块,显著提升检测准确率
- 可对接大语言模型,增强结果可解释性
在当前数字环境中,多模态媒体伪造已成为影响信息可信度的重大威胁。现有检测方法常忽视局部信息,而伪造往往集中在面部区域。为此,我们提出M4-BLIP框架,以广受认可的BLIP-2模型为基础提取局部特征,并引入人脸先验知识。框架内设计了专门的对齐与融合模块,精细整合局部与全局特征,实现更优的检测性能。此外,该方法可无缝接入大型语言模型(LLM),显著提升检测结果的可解释性。大量定量与可视化实验验证了其在主流竞争方法上的优越性。
原文摘要 · Abstract (English)
In the contemporary digital landscape, multi-modal media manipulation has emerged as a significant societal threat, impacting the reliability and integrity of information dissemination. Current detection methodologies in this domain often overlook the crucial aspect of localized information, despite the fact that manipulations frequently occur in specific areas, particularly in facial regions. In response to this critical observation, we propose the M4-BLIP framework. This innovative framework utilizes the BLIP-2 model, renowned for its ability to extract local features, as the cornerstone for feature extraction. Complementing this, we incorporate local facial information as prior knowledge. A specially designed alignment and fusion module within M4-BLIP meticulously integrates these local and global features, creating a harmonious blend that enhances detection accuracy. Furthermore, our approach seamlessly integrates with Large Language Models (LLM), significantly improving the interpretability of the detection outcomes. Extensive quantitative and visualization experiments validate the effectiveness of our framework against the state-of-the-art competitors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。