arXiv:2604.27218cs.CV2026-04

量化人体嵌入中属性表达强度,揭示模型隐含偏见与跨模态差异。

AttriBE: Quantifying Attribute Expressivity in Body Embeddings for Recognition and Identification

论文配图:AttriBE: Quantifying Attribute Expressivity in Body Embeddings for Recognition and Identification
图 1 · 摘自论文原文
  • 用二级神经网络测量特征与性别、姿态等属性的互信息。
  • 深度层中体型指数(BMI)表达最强,姿态在中间层峰值出现。
  • 红外跨模态识别时结构线索更关键,适合关注公平性与鲁棒性的研究者。

行人重识别(ReID)系统在多图或视频帧间匹配个体至关重要,但现有方法常受性别、姿态和身体质量指数(BMI)等属性干扰,尤其在非约束场景下影响公平性与泛化能力。为此,我们扩展了表达性概念——即学习特征与特定属性间的互信息,并通过二级神经网络量化属性编码强度。在大规模可见光数据集上对三种基于Transformer的ReID模型进行分析发现,BMI在深层特征中表达性始终最高,属性表达强度排序为:BMI > Pitch > Gender > Yaw;且表达性随网络层级与训练轮次演化,姿态在中间层达峰值,而BMI随深度增强。进一步扩展至短波、中波与长波红外跨谱行人识别,发现此时俯仰角表达性接近BMI,属性趋势沿深度单调上升,表明跨模态时对结构线索依赖更强。总体显示,Transformer-based ReID嵌入中存在隐式属性层次结构,体型信息持续保留,姿态在跨谱条件下贡献更大。

原文摘要 · Abstract (English)

Person re-identification (ReID) systems that match individuals across images or video frames are essential in many real-world applications. However, existing methods are often influenced by attributes such as gender, pose, and body mass index (BMI), which vary in unconstrained settings and raise concerns related to fairness and generalization. To address this, we extend the notion of expressivity, defined as the mutual information between learned features and specific attributes, using a secondary neural network to quantify how strongly attributes are encoded. Applying this framework to three transformer-based ReID models on a large-scale visible-spectrum dataset, we find that BMI consistently shows the highest expressivity in deeper layers. Attributes in the final representation are ranked as BMI > Pitch > Gender > Yaw, and expressivity evolves across layers and training epochs, with pose peaking in intermediate layers and BMI strengthening with depth. We further extend the analysis to cross-spectral person identification across infrared modalities including short-wave, medium-wave, and long-wave infrared. In this setting, pitch becomes comparable to BMI and attribute trends increase monotonically across depth, suggesting increased reliance on structural cues when bridging modality gaps. Overall, the results show that transformer-based ReID embeddings encode a hierarchy of implicit attributes, with morphometric information persistently embedded and pose contributing more strongly under cross-spectral conditions.

行人重识别属性表达跨模态公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。