arXiv:2506.12447cs.CV2025-06被引 3

用手部图像实现精准身份识别,助力性侵等案件破案

CLIP-HandID: Vision-Language Model for Hand-Based Person Identification

  • 基于CLIP模型,用文本提示引导手部特征学习
  • 通过伪标记技术生成视觉语义,提升识别准确率
  • 适用于无面部信息的犯罪调查,适合刑侦人员使用

本文提出一种基于手部图像的人体识别新方法——CLIP-HandID,专为刑事案件中缺乏面部信息的场景设计。在性侵等严重犯罪中,手部图像常是唯一可获取的生物证据。该方法利用预训练的视觉语言模型CLIP,通过文本提示作为语义引导,从手部图像中提取判别性特征表示。由于手部图像仅以编号标注而非文字描述,研究引入文本反演网络,学习能编码特定视觉上下文或外观属性的伪标记,并将其融入文本提示输入到CLIP的文本编码器中,以激发多模态推理能力,增强模型泛化性能。在两个具有多元种族代表性的大型公开手部数据集上进行的大量实验表明,该方法显著优于现有方法。

原文摘要 · Abstract (English)

This paper introduces a novel approach to person identification using hand images, designed specifically for criminal investigations. The method is particularly valuable in serious crimes such as sexual abuse, where hand images are often the only identifiable evidence available. Our proposed method, CLIP-HandID, leverages a pre-trained foundational vision-language model - CLIP - to efficiently learn discriminative deep feature representations from hand images (input to CLIP's image encoder) using textual prompts as semantic guidance. Since hand images are labeled with indexes rather than text descriptions, we employ a textual inversion network to learn pseudo-tokens that encode specific visual contexts or appearance attributes. These learned pseudo-tokens are then incorporated into textual prompts, which are fed into CLIP's text encoder to leverage its multi-modal reasoning and enhance generalization for identification. Through extensive evaluations on two large, publicly available hand datasets with multi-ethnic representation, we demonstrate that our method significantly outperforms existing approaches.

手部识别多模态刑侦应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。