arXiv:2507.01290cs.CV2025-07

用知识令牌高效融合多模型特征,提升人脸分析精度

Learning an Ensemble Token from Task-driven Priors in Facial Analysis

  • 设计知识令牌,在自注意力中统一预训练编码器的互信息
  • 在多个任务上实现显著性能提升,计算开销几乎为零
  • 适合需要轻量高精度人脸分析的场景

人脸分析存在任务特异性特征差异。尽管卷积神经网络(CNN)能精细表征空间信息,视觉变换器(ViT)可在块级别捕捉语义信息,但过去十年骨干网络的进步常伴随特征表示层面的高计算成本。本文提出KT-Adapter,一种学习知识令牌的新方法,以高效整合高保真特征表示。具体地,我们设计了鲁棒先验统一学习方法,在自注意力机制内生成知识令牌,共享多个预训练编码器间的互信息。该知识令牌方法具有极高效率,计算开销可忽略不计。实验表明,该方法在各类人脸分析任务中均取得显著性能提升,特征表示能力得到统计学意义上的增强。

原文摘要 · Abstract (English)

Facial analysis exhibits task-specific feature variations. While Convolutional Neural Networks (CNNs) have enabled the fine-grained representation of spatial information, Vision Transformers (ViTs) have facilitated the representation of semantic information at the patch level. While advances in backbone architectures have improved over the past decade, combining high-fidelity models often incurs computational costs on feature representation perspective. In this work, we introduce KT-Adapter, a novel methodology for learning knowledge token which enables the integration of high-fidelity feature representation in computationally efficient manner. Specifically, we propose a robust prior unification learning method that generates a knowledge token within a self-attention mechanism, sharing the mutual information across the pre-trained encoders. This knowledge token approach offers high efficiency with negligible computational cost. Our results show improved performance across facial analysis, with statistically significant enhancements observed in the feature representations.

人脸分析知识令牌高效模型ViT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。