用思维链提升多模态模型防人脸欺骗能力
Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing
- 引入视觉语言联合推理机制,增强模型理解力
- 在14类攻击上性能超越现有方法,提升显著
- 适合需要高鲁棒性与可解释性的安全场景
人脸反欺骗(FAS)传统上依赖单一视觉模态,难以应对打印、屏幕重放、3D面具等多样攻击,泛化能力受限。多模态大语言模型(MLLM)在图像-文本理解与语义推理方面取得突破,提示将视觉与语言协同推理引入FAS可显著提升鲁棒性与可解释性。然而,高质量视觉-语言多模态数据集匮乏成为关键瓶颈。为此,我们提出首个面向FAS的大型视觉问答(VQA)数据集FaceCoT,涵盖14类欺骗攻击,并提供高质量思维链(CoT)标注。同时,采用强化学习微调的描述生成模型扩展数据并提升标注质量。此外,设计了基于思维链的渐进式学习策略(CEPL),有效利用CoT数据,显著提升模型在FAS任务上的表现。大量实验表明,使用FaceCoT与CEPL训练的模型在多个基准数据集上均优于当前最先进方法。
原文摘要 · Abstract (English)
Face Anti-Spoofing (FAS) typically depends on a single visual modality when defending against presentation attacks such as print attacks, screen replays, and 3D masks, resulting in limited generalization across devices, environments, and attack types. Meanwhile, Multimodal Large Language Models (MLLMs) have recently achieved breakthroughs in image-text understanding and semantic reasoning, suggesting that integrating visual and linguistic co-inference into FAS can substantially improve both robustness and interpretability. However, the lack of a high-quality vision-language multimodal dataset has been a critical bottleneck. To address this, we introduce FaceCoT (Face Chain-of-Thought), the first large-scale Visual Question Answering (VQA) dataset tailored for FAS. FaceCoT covers 14 spoofing attack types and enriches model learning with high-quality CoT VQA annotations. Meanwhile, we develop a caption model refined via reinforcement learning to expand the dataset and enhance annotation quality. Furthermore, we introduce a CoT-Enhanced Progressive Learning (CEPL) strategy to better leverage the CoT data and boost model performance on FAS tasks. Extensive experiments demonstrate that models trained with FaceCoT and CEPL outperform state-of-the-art methods on multiple benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。