用语言提示+双路径模型,让远距离识性别更准更可靠。
When Gender is Hard to See: Multi-Attribute Support for Long-Range Recognition
- 双路径设计:直接视觉+语言提示引导的属性推理
- 在长距离下多指标超越现有方法,对视角变化鲁棒
- 适合需要可解释性与责任拒答的远距离识别场景
由于空间分辨率低、视角变化大、面部线索缺失,极远距离图像中的性别识别仍具挑战。为此,我们提出一种基于CLIP的双路径Transformer框架,联合建模视觉与属性驱动线索。该框架包含两个互补分支:(1)直接视觉路径,通过选择性微调CLIP图像编码器高层层进行优化;(2)属性引导路径,利用发型、服装、配饰等软生物特征提示,在CLIP文本-图像空间中推断性别。空间通道注意力模块进一步增强遮挡和低分辨率下的判别定位能力。为支持大规模评估,我们构建了U-DetAGReID数据集,源自DetReIDx和AG-ReID.v2,采用统一的三元标签体系(男、女、未知)。大量实验表明,所提方法在宏平均F1、准确率、AUC等多个指标上优于当前最优的人体属性识别与重识别基线,对距离、角度、高度变化保持一致鲁棒性。定性注意力可视化证实了属性定位的可解释性与负责任的拒答行为。结果表明,语言引导的双路径学习为开放场景下的负责任性别识别提供了可扩展的理论基础。
原文摘要 · Abstract (English)
Accurate gender recognition from extreme long-range imagery remains a challenging problem due to limited spatial resolution, viewpoint variability, and loss of facial cues. For such purpose, we present a dual-path transformer framework that leverages CLIP to jointly model visual and attribute-driven cues for gender recognition at a distance. The framework integrates two complementary streams: (1) a direct visual path that refines a pre-trained CLIP image encoder through selective fine-tuning of its upper layers, and (2) an attribute-mediated path that infers gender from a set of soft-biometric prompts (e.g., hairstyle, clothing, accessories) aligned in the CLIP text-image space. Spatial channel attention modules further enhance discriminative localization under occlusion and low resolution. To support large-scale evaluation, we construct U-DetAGReID, a unified long-range gender dataset derived from DetReIDx and AG-ReID.v2, harmonized under a consistent ternary labeling scheme (Male, Female, Unknown). Extensive experiments suggest that the proposed solution surpasses state-of-the-art person-attribute and re-identification baselines across multiple metrics (macro-F1, accuracy, AUC), with consistent robustness to distance, angle, and height variations. Qualitative attention visualizations confirm interpretable attribute localization and responsible abstention behavior. Our results show that language-guided dual-path learning offers a principled, extensible foundation for responsible gender recognition in unconstrained long-range scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。