arXiv:2607.08156cs.CV2026-07中稿 · ICME 2026

通过细粒度文本描述提升人脸识别攻击检测的泛化能力

Unified Face Attack Detection via Fine-Grained Semantic Guidance

论文配图:Unified Face Attack Detection via Fine-Grained Semantic Guidance
图 1 · 摘自论文原文
  • 用细粒度文本描述伪造线索,指导模型理解攻击本质
  • 在800万张攻击图像上验证,性能超越纯视觉与粗粒度描述方法
  • 适合关注安全检测与多模态融合的研究者

人脸识别系统广泛应用的同时面临日益多样化的安全威胁。现有数据集缺乏对伪造线索的详细文本描述,导致多数先前方法将人脸攻击检测视为纯视觉识别任务。本文基于包含超过800万张攻击图像的大规模MS-UFAD数据集,为每张图像添加细粒度的伪造线索文本描述,并提出双对齐伪造网络(DAF-Net)以更有效地利用这些文本信息。大量实验表明,该方法能从攻击图像中提取更具泛化性和语义意义的伪造表征,优于仅依赖视觉或使用粗粒度描述的方法。

原文摘要 · Abstract (English)

The growing applications of facial recognition systems are accompanied by increasingly diverse security threats. Existing datasets lack detailed textual descriptions of forgery cues, leading most prior methods to treat face attack detection primarily as a visual recognition task. In this paper, building upon the large-scale MS-UFAD dataset which contains over 8 million attack images, we enrich each image with a fine-grained textual description of forgery cues. Furthermore, we propose a Dual Alignment Forgery Network(DAF-Net) to better leverage these textual information. Extensive experiments demonstrate that our approach extracts more generalizable and semantically meaningful forgery representations from attack images, outperforming both vision-only methods and approaches based on coarse-grained descriptions.

人脸识别攻击检测多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。