arXiv:2608.09633cs.CVcs.LG2026-08

LoRA微调对人脸活体攻击检测效果有限,跨数据集泛化能力差。

LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection

论文配图:LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection
图 1 · 摘自论文原文
  • 用32种基础模型对比评估,仅靠低秩适配(LoRA)无法提升泛化性能。
  • 跨数据集时错误率超2%,远高于单数据集下的近0%错误率。
  • 适合关注模型泛化能力与轻量适配局限的研究者阅读。

人脸活体攻击检测(PAD)旨在可靠识别各类攻击手段。尽管现有方法在单一数据集上表现优异,但在跨数据集测试中性能显著下降,传感器或光照变化可使检测准确率从近乎完美降至接近随机。基础模型(FMs)因其预训练规模远超传统小样本PAD数据集(如MCIO基准:MSU-MFSD、CASIA-FASD、Replay-Attack、OULU-NPU),被视为有前景的解决方案。然而,当前研究多集中于CLIP类模型,忽视了其他架构与训练方式的基础模型。本研究系统评估了32种基础模型。零样本提示在各模型族与规模下均接近随机水平。视觉编码器经LoRA低秩适配(参数量少于1%)后,在多数情况下单数据集ACER低于2%,但跨数据集ACER显著升高。表明LoRA主要优化数据集内判别边界,模型预训练表征与适配数据集对跨域泛化的影响大于轻量适配策略本身。

原文摘要 · Abstract (English)

Face presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. While PAD methods achieve strong performance within individual datasets, their performance degrades under cross-dataset evaluation. Variations in sensors or lighting conditions can reduce the effectiveness of detectors from near-perfect to nearly random. Foundation models (FMs) have emerged as a promising alternative because typical PAD datasets, such as the MCIO benchmarks (MSU-MFSD, CASIA-FASD, Replay-Attack, and OULU-NPU), are small relative to the scale used for web-based pretraining. However, existing PAD systems primarily focus on CLIP-based foundation models, while overlooking other FMs with different architectures and training procedures. This study addresses this question by systematically evaluating 32 FMs. Zero-shot prompting achieves performance near chance across model families and scales. The vision encoders, when low-rankadapted (LoRA) with fewer than 1% trainable weights, achieve below 2% intra-dataset ACER in most cases, while cross-dataset ACER is substantially higher. LoRA primarily refines the decision boundary within a dataset, suggesting that pretrained representations and the adaptation dataset play a larger role in cross-dataset generalization than the evaluated lightweight adaptation strategy.

人脸检测基础模型泛化能力LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。