用少量高质量标注训练可解释的面部反欺骗系统,效果远超传统方法。
MS-MFAD : Multimodal large language models for Face Anti-spoofing Detection

- 通过像素级语义锚定激活多模态大模型内在推理能力
- 仅用1000张精准掩码实现40%-50%的误报率降低
- 支持可审计推理路径,适合需要可信验证的安防场景
当前人脸识别系统面临生成式AI与高保真物理伪造的复合威胁。现有防御方法存在泛化差、推理不可审计、依赖海量低质数据等瓶颈。为此,我们提出基于多模态大语言模型的面部反欺骗检测框架(MFAD),构建统一的面部反欺骗检测(UFAD)可解释推理系统,并建立语义级标注基准。不同于依赖外部工具或粗粒度对齐的方法,MFAD通过细粒度像素-语义锚定机制激活多模态大语言模型(MLLMs)的内在推理能力,消除定位幻觉,确保推理路径可审计。我们提出跨攻击类型的语义级统一标注范式:仅需每类攻击标注1000张精确掩码,即可生成严格对应伪造区域的推理证据链。在Qwen-VL基础模型上进行监督微调,仅用少量高质量样本即实现域内ACER降低40%-50%,跨域性能下降控制在11.62%/5.23%以内,显著优于现有框架。白盒对抗攻击下,检测准确率仅下降3.2%,验证了语义锚定的鲁棒性。领域专家对推理路径可靠性评分达4.57/5,推理延迟满足实时部署要求。结果表明,少样本高质语义标注范式能有效构建可信、可解释且低成本的UFAD系统。
原文摘要 · Abstract (English)
Facial biometric recognition systems currently face compound threats intertwining generative AI and high-fidelity physical spoofing. Existing defenses suffer from systemic bottlenecks, including poor generalization, non-auditable reasoning, and reliance on massive, low-quality datasets. To address these challenges, we propose Multimodal Large Language Models (MFAD) for face anti-spoofing detection, an explainable reasoning system for Unified Face Anti-Spoofing Detection (UFAD), accompanied by a semantic-level annotation benchmark. Unlike methods relying on external tools or coarse alignment, MFAD activates the intrinsic reasoning capabilities of Multimodal Large Language Models (MLLMs) via a fine-grained pixel-semantic anchoring mechanism. This eliminates localization hallucinations and ensures auditable reasoning paths. We introduce a cross-attack semantic-level unified annotation paradigm: by annotating only 1,000 precise masks per attack category, we generate reasoning evidence chains strictly corresponding to spoofed regions. Supervised fine-tuning on the Qwen-VL foundation model demonstrates that, using limited high-quality samples, the system achieves a 40-50% relative reduction in in-domain ACER and restricts cross-domain performance degradation to within 11.62%/5.23%, significantly outperforming existing frameworks. Furthermore, under white-box adversarial attacks, detection accuracy drops by only 3.2%, validating the robustness of semantic anchoring compared to models trained on massive short-text data. Domain practitioners rated the evidence reliability of reasoning paths at 4.57/5, with inference latency satisfying real-time deployment requirements. These results confirm that a few-shot, high-quality semantic annotation paradigm is effective for building trustworthy, explainable, and cost-efficient UFAD systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。