通过增强推理路径提升多模态人脸反欺骗的可解释性与泛化能力
PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
- 构建扩展推理链,利用有限标注生成高质量多模态推理路径
- 引入答案洗牌机制,避免模型依赖表面线索,提升跨模态分析深度
- 显著提升跨域泛化性能,实现融合、泛化与可解释性的统一
多模态人脸反欺骗(FAS)在跨域泛化和可解释性方面取得进展。结合大语言模型与强化学习(RL),基于策略的训练为联合建模这些特性提供了新机遇。然而,多模态推理比单模态更复杂,需精准特征表示与跨模态验证,且面临标注稀缺、质量低的问题,直接应用RL效果不佳。我们发现监督微调加强化学习(SFT+RL)在多模态FAS中的两个关键局限:(1) 推理路径受限,限制互补模态利用,使探索空间缩小;(2) 单任务监督与多样化推理路径不匹配,导致推理混淆,模型可能绕过真实推理过程,仅通过图像直接映射答案。为此,我们提出PA-FAS,通过从有限标注中构建高质量扩展推理序列,丰富推理路径并放松探索约束。同时在微调阶段引入答案洗牌机制,强制进行全面多模态分析,抑制捷径学习。PA-FAS显著提升多模态推理准确率与跨域泛化能力,更好统一多模态融合、泛化与可解释性,推动可信FAS发展。
原文摘要 · Abstract (English)
Face anti-spoofing (FAS) has recently advanced in multimodal fusion, cross-domain generalization, and interpretability. With large language models and reinforcement learning (RL), strategy-based training offers new opportunities to jointly model these aspects. However, multimodal reasoning is more complex than unimodal reasoning, requiring accurate feature representation and cross-modal verification while facing scarce, high-quality annotations, which makes direct application of RL sub-optimal. We identify two key limitations of supervised fine-tuning plus RL (SFT+RL) for multimodal FAS: (1) limited multimodal reasoning paths restrict the use of complementary modalities and shrink the exploration space after SFT, weakening the effect of RL; and (2) mismatched single-task supervision versus diverse reasoning paths causes reasoning confusion, where models may exploit shortcuts by mapping images directly to answers and ignoring the intended reasoning. To address this, we propose PA-FAS, which enhances reasoning paths by constructing high-quality extended reasoning sequences from limited annotations, enriching paths and relaxing exploration constraints. We further introduce an answer-shuffling mechanism during SFT to force comprehensive multimodal analysis instead of using superficial cues, thereby encouraging deeper reasoning and mitigating shortcut learning. PA-FAS significantly improves multimodal reasoning accuracy and cross-domain generalization, and better unifies multimodal fusion, generalization, and interpretability for trustworthy FAS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。