arXiv:2511.11754cs.CV2025-11

提出新型批量注意力架构,提升表情识别合成图像生成效率

Batch Transformer Architecture: Case of Synthetic Image Generation for Emotion Expression Facial Recognition

  • 仅关注重要特征维度,降低编码器-解码器瓶颈
  • 在妆容与遮挡数据集上实现有限数据的高变异性生成
  • 适合需要高效生成多样化人脸图像的研究者

提出一种隐式稀疏风格的新型Transformer变体。与传统Transformer不同,该架构不关注整个维度序列或批次实体,而是聚焦于关键特征维度(主成分)。通过选择重要维度进行注意力计算,显著减小了编码器-解码器神经网络架构中的瓶颈规模。该架构在妆容与遮挡数据集上测试,用于人脸表情识别的合成图像生成,有效提升了原始数据集的多样性。

原文摘要 · Abstract (English)

A novel Transformer variation architecture is proposed in the implicit sparse style. Unlike "traditional" Transformers, instead of attention to sequential or batch entities in their entirety of whole dimensionality, in the proposed Batch Transformers, attention to the "important" dimensions (primary components) is implemented. In such a way, the "important" dimensions or feature selection allows for a significant reduction of the bottleneck size in the encoder-decoder ANN architectures. The proposed architecture is tested on the synthetic image generation for the face recognition task in the case of the makeup and occlusion data set, allowing for increased variability of the limited original data set.

Transformer图像生成特征选择人脸识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。