用工具增强大模型推理,提升人脸反欺骗的泛化能力
From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing
- 将人脸识别反欺骗任务转为带视觉工具的思维链推理
- 在11个域间测试中达到最新最佳性能,识别细微伪造痕迹
- 适合关注可解释性与鲁棒性的安全检测研究者
人脸识别仍易受呈现攻击,亟需鲁棒的面部反欺骗(FAS)方案。现有基于多模态大模型(MLLM)的方法将二分类任务转化为生成简短文本描述,以提升跨域泛化能力,但其泛化性能受限,因描述主要捕捉直观语义线索(如口罩轮廓),难以察觉细微视觉模式。为此,我们引入外部视觉工具增强MLLM,推动对微弱欺骗线索的深度探查。提出工具增强推理框架TAR-FAS,将FAS任务重构为带视觉工具的思维链(CoT-VT)范式,使模型从直观观察出发,动态调用外部工具进行细粒度分析。构建包含多轮工具使用推理轨迹的ToolFAS-16K数据集,并设计工具感知训练流程,其中多样工具组相对策略优化(DT-GRPO)使模型自主学习高效工具使用。在挑战性的一对十一跨域协议下实验表明,TAR-FAS达到当前最优性能,同时提供可信的细粒度视觉探查过程。
原文摘要 · Abstract (English)
Face recognition remains vulnerable to presentation attacks, calling for robust Face Anti-Spoofing (FAS) solutions. Recent MLLM-based FAS methods reformulate the binary classification task as the generation of brief textual descriptions to improve cross-domain generalization. However, their generalizability is still limited, as such descriptions mainly capture intuitive semantic cues (e.g., mask contours) while struggling to perceive fine-grained visual patterns. To address this limitation, we incorporate external visual tools into MLLMs to encourage deeper investigation of subtle spoof clues. Specifically, we propose the Tool-Augmented Reasoning FAS (TAR-FAS) framework, which reformulates the FAS task as a Chain-of-Thought with Visual Tools (CoT-VT) paradigm, allowing MLLMs to begin with intuitive observations and adaptively invoke external visual tools for fine-grained investigation. To this end, we design a tool-augmented data annotation pipeline and construct the ToolFAS-16K dataset, which contains multi-turn tool-use reasoning trajectories. Furthermore, we introduce a tool-aware FAS training pipeline, where Diverse-Tool Group Relative Policy Optimization (DT-GRPO) enables the model to autonomously learn efficient tool use. Extensive experiments under a challenging one-to-eleven cross-domain protocol demonstrate that TAR-FAS achieves SOTA performance while providing fine-grained visual investigation for trustworthy spoof detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。