arXiv:2608.21455cs.CV2026-08中稿 · ACM Multimedia 202…

用番茄土豆洋葱做攻击检测,发现无需人脸也能学出有效防御方法。

Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection

论文配图:Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection
图 1 · 摘自论文原文
  • 用果蔬代替人脸训练攻击检测模型,验证无脸可行性
  • 无脸模型跨数据集平均AUC达92.70%,优于合成人脸训练
  • 适合隐私保护和不依赖身份的安防系统研发

传统人脸活体检测(PAD)聚焦于面部特征,但打印、重放等攻击引入的视觉痕迹并不专属人脸。本文提出TPO数据集,包含番茄、土豆、洋葱在类人脸PAD协议下的真实、打印和重放样本。基于基础模型架构,仅用TPO训练的检测器在四个标准跨数据集人脸PAD基准上平均AUC达92.70%,超越合成人脸训练,且媲美真实人脸训练模型。反向测试显示,人脸模型在TPO上表现持续高于随机水平,说明其学习的是攻击过程特征而非物体语义。将TPO加入常规人脸训练可稳定提升跨数据集性能,在固定优化预算下提供互补信息。频域与表示分析表明,可迁移特征并非单一频谱伪影,而是跨类别共享的丰富攻击线索。结果证明:可迁移的攻击检测表征可脱离人脸内容独立学习,为隐私保护与身份无关的PAD发展开辟新路径。

原文摘要 · Abstract (English)

Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.

活体检测无脸训练隐私保护跨数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。