构建首个统一人脸伪造检测数据集,实现物理与数字攻击的联合识别。
Benchmarking Unified Face Attack Detection via Hierarchical Prompt Tuning
- 提出分层提示调优框架,自适应探索多语义空间分类规则。
- 构建包含69万视频的UniAttackData+数据集,覆盖54类伪造技术。
- 适合需要跨类型伪造检测的安防与内容可信系统研发者。
PAD和FFD分别用于防护基于实体介质的演示攻击和基于数字编辑的DeepFakes,但两者独立训练会增加对未知攻击的脆弱性,给部署环境带来负担。缺乏能同时处理这两类攻击的统一人脸伪造检测模型,主要源于两点:(1) 缺乏足够丰富的基准数据集。现有UAD数据集仅包含有限攻击类型与样本,导致模型应对复杂威胁能力受限。为此,我们通过可解释的分层方式,构建迄今最全面、最复杂的伪造技术集合——UniAttackDataPlus,涵盖2,875个身份及其对应的54类伪造样本,共计697,347个视频。(2) 缺乏可信的分类标准。现有方法试图在相同语义空间内探索任意准则,但在面对多样化攻击时难以成立。因此,我们提出基于视觉-语言模型的分层提示调优框架,可自适应从不同语义空间中探索多种分类准则。具体而言,构建VP-Tree以分层探索各类分类规则;通过自适应剪枝提示,模型可引导编码器在粗到细的不同层级提取判别性特征。最后,为帮助模型理解视觉空间的分类标准,提出DPI模块,将视觉提示投影至文本编码器,以获得更准确的语义表示。
原文摘要 · Abstract (English)
PAD and FFD are proposed to protect face data from physical media-based Presentation Attacks and digital editing-based DeepFakes, respectively. However, isolated training of these two models significantly increases vulnerability towards unknown attacks, burdening deployment environments. The lack of a Unified Face Attack Detection model to simultaneously handle attacks in these two categories is mainly attributed to two factors: (1) A benchmark that is sufficient for models to explore is lacking. Existing UAD datasets only contain limited attack types and samples, leading to the model's confined ability to address abundant advanced threats. In light of these, through an explainable hierarchical way, we propose the most extensive and sophisticated collection of forgery techniques available to date, namely UniAttackDataPlus. Our UniAttackData+ encompasses 2,875 identities and their 54 kinds of corresponding falsified samples, in a total of 697,347 videos. (2) The absence of a trustworthy classification criterion. Current methods endeavor to explore an arbitrary criterion within the same semantic space, which fails to exist when encountering diverse attacks. Thus, we present a novel Visual-Language Model-based Hierarchical Prompt Tuning Framework that adaptively explores multiple classification criteria from different semantic spaces. Specifically, we construct a VP-Tree to explore various classification rules hierarchically. Then, by adaptively pruning the prompts, the model can select the most suitable prompts guiding the encoder to extract discriminative features at different levels in a coarse-to-fine manner. Finally, to help the model understand the classification criteria in visual space, we propose a DPI module to project the visual prompts to the text encoder to help obtain a more accurate semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。