发现人脸模板中的语义概念,实现无需重编码的可控编辑。
SCOUT: Semantic Concept Discovery for Open-Vocabulary Editing of face Recognition Templates

- 通过自然语言生成语义假设,自动发现模板中的可解释概念。
- 在多种模型上成功识别出超越标准属性的新语义方向。
- 适合需要精准控制人脸识别模板的开发者和研究者使用。
人脸识别模板是紧凑的身份表征,同时包含丰富的外观语义信息。以往工作可通过图像反演或间接编辑管线操作模板,但直接在模板空间进行语义编辑仍基本未被探索。现有可解释性方法依赖人工神经元分析或预定义属性标签,难以扩展且语义灵活性不足。为此,我们提出SCOUT(面向开放词汇的人脸识别模板语义概念发现),一个端到端框架,利用机制可解释性在模板空间中发现并直接操纵语义概念。SCOUT学习稀疏模板表示,从自然语言描述生成潜在特征语义假设,并验证其稳定性。所得特征作为可调控的语义方向,实现无需代价高昂的编辑-重编码流程的直接编辑。在采用CNN、ViT和Swin主干网络的人脸识别模型上实验表明,SCOUT发现了超出标准属性标签的可解释概念,实现了可控且身份感知的模板操控,对身份匹配性能影响极小。此外,经编辑的模板可借助独立反演模型解码用于可视化与评估。
原文摘要 · Abstract (English)
Face recognition templates are compact identity representations, yet they also encode rich semantic information about facial appearance. Prior work has shown that templates can be inverted to images or indirectly manipulated through image-editing pipelines, but direct semantic editing in template space remains largely unexplored. Existing interpretability methods for face recognition often rely on manual neuron inspection or predefined attribute labels, limiting scalability and semantic flexibility. To address this gap, we propose SCOUT (Semantic Concept Discovery for Open-VocabUlary Editing of Face Recognition Templates), an end-to-end framework for discovering and directly manipulating semantic concepts in face recognition templates using mechanistic interpretability. SCOUT learns sparse template representations, generates semantic hypotheses for latent features from natural-language descriptions, and validates their stability. The resulting features act as controllable semantic directions for direct editing, avoiding costly edit--re-encode pipelines. Experiments with face recognition models using CNN, ViT, and Swin backbones show that SCOUT discovers interpretable concepts beyond standard attribute labels and enables controllable, identity-aware template manipulation with negligible impact on identity matching. We further show that edited templates can subsequently be decoded with independent inversion models for visualization and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。