用视觉语言模型实现无数据重演的持续抗欺骗检测。
Steering Vision-Language Pre-trained Models for Incremental Face Presentation Attack Detection
- 通过多方面提示与选择性权重保持,平衡新旧任务学习。
- 在多个基准上显著减少灾难性遗忘,提升未知场景性能。
- 适合需隐私保护的长期人脸识别系统部署。
人脸展示攻击检测(PAD)需持续学习以应对不断演变的伪造手段和场景。然而,隐私法规禁止保留历史数据,因此要求无重演持续学习(RF-IL)。视觉语言预训练(VLP)模型具备可调提示的跨模态表征,能高效适应新伪造风格与域。本文提出基于VLP的RF-IL框架SVLP-IL,通过多方面提示(MAP)和选择性弹性权重保持(SEWC)实现稳定与灵活性的平衡。MAP分离领域依赖,增强对分布偏移的敏感性,并通过通用与领域特定线索联合缓解遗忘;SEWC选择性保留前任务关键权重,保留知识的同时支持新任务适应。在多个PAD基准上的实验表明,SVLP-IL显著降低灾难性遗忘,提升对未见域的性能。该方法为隐私合规的鲁棒终身PAD部署提供了实用解决方案。
原文摘要 · Abstract (English)
Face Presentation Attack Detection (PAD) demands incremental learning (IL) to combat evolving spoofing tactics and domains. Privacy regulations, however, forbid retaining past data, necessitating rehearsal-free IL (RF-IL). Vision-Language Pre-trained (VLP) models, with their prompt-tunable cross-modal representations, enable efficient adaptation to new spoofing styles and domains. Capitalizing on this strength, we propose \textbf{SVLP-IL}, a VLP-based RF-IL framework that balances stability and plasticity via \textit{Multi-Aspect Prompting} (MAP) and \textit{Selective Elastic Weight Consolidation} (SEWC). MAP isolates domain dependencies, enhances distribution-shift sensitivity, and mitigates forgetting by jointly exploiting universal and domain-specific cues. SEWC selectively preserves critical weights from previous tasks, retaining essential knowledge while allowing flexibility for new adaptations. Comprehensive experiments across multiple PAD benchmarks show that SVLP-IL significantly reduces catastrophic forgetting and enhances performance on unseen domains. SVLP-IL offers a privacy-compliant, practical solution for robust lifelong PAD deployment in RF-IL settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。