arXiv:2505.09484cs.CV2025-05中稿 · ECCV被引 3

提出MMDA框架,通过去噪与对齐提升多模态人脸反伪造的跨域泛化能力。

Purify then Guide: Rethinking Domain Generalization for Multimodal Face Anti-Spoofing

  • 设计联合噪声注意力模块,同时抑制模态与域间干扰
  • 在4个数据集上实现最优跨域检测准确率,超越现有方法
  • 适合关注跨场景人脸识别安全的研究者与工程师

人脸反伪造(FAS)对支付和监控等场景的面部识别安全至关重要。现有多模态FAS方法常因模态特异性偏差和域偏移导致泛化能力不足。本文提出多模态去噪与对齐(MMDA)框架,利用CLIP的零样本泛化能力,通过去噪与对齐机制有效抑制多模态数据噪声,显著提升跨模态对齐的泛化性能。其中,模态-域联合差分注意力(MD2A)模块基于提取的共性噪声特征优化注意力机制,同时缓解域与模态噪声影响。代表性空间软对齐(RS2)策略借助预训练CLIP模型,灵活地将多域多模态数据对齐至通用表示空间,保留复杂表征并增强模型对未见条件的适应性。此外,设计了U形双空间自适应(U-DSA)模块,在保持泛化能力的同时提升表示适应性。实验结果表明,MMDA在四个基准数据集的不同评估协议下,均优于当前最先进方法,显著提升跨域泛化与多模态检测精度。代码即将开源。

原文摘要 · Abstract (English)

Face Anti-Spoofing (FAS) is essential for the security of facial recognition systems in diverse scenarios such as payment processing and surveillance. Current multimodal FAS methods often struggle with effective generalization, mainly due to modality-specific biases and domain shifts. To address these challenges, we introduce the \textbf{M}ulti\textbf{m}odal \textbf{D}enoising and \textbf{A}lignment (\textbf{MMDA}) framework. By leveraging the zero-shot generalization capability of CLIP, the MMDA framework effectively suppresses noise in multimodal data through denoising and alignment mechanisms, thereby significantly enhancing the generalization performance of cross-modal alignment. The \textbf{M}odality-\textbf{D}omain Joint \textbf{D}ifferential \textbf{A}ttention (\textbf{MD2A}) module in MMDA concurrently mitigates the impacts of domain and modality noise by refining the attention mechanism based on extracted common noise features. Furthermore, the \textbf{R}epresentation \textbf{S}pace \textbf{S}oft (\textbf{RS2}) Alignment strategy utilizes the pre-trained CLIP model to align multi-domain multimodal data into a generalized representation space in a flexible manner, preserving intricate representations and enhancing the model's adaptability to various unseen conditions. We also design a \textbf{U}-shaped \textbf{D}ual \textbf{S}pace \textbf{A}daptation (\textbf{U-DSA}) module to enhance the adaptability of representations while maintaining generalization performance. These improvements not only enhance the framework's generalization capabilities but also boost its ability to represent complex representations. Our experimental results on four benchmark datasets under different evaluation protocols demonstrate that the MMDA framework outperforms existing state-of-the-art methods in terms of cross-domain generalization and multimodal detection accuracy. The code will be released soon.

人脸反伪造多模态域泛化CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。