arXiv:2507.01504cs.CVcs.AI2025-07中稿 · publication at the…被引 1

用跨模态线索提升行人重识别,更好保护隐私

Following the Clues: Experiments on Person Re-ID using Cross-Modal Intelligence

  • 融合视觉语言模型与图注意力网络,识别文本可描述的隐私线索
  • 在Market-1501到CUHK03-np跨数据集测试中表现更优
  • 适合关注隐私保护与可解释性重识别的研究者

街景视频数据的开放对自动驾驶和AI研究至关重要,但其中包含的个人身份信息(PII)远超面部等生物特征,带来严重隐私风险。本文提出cRID框架,结合大视觉语言模型、图注意力网络与表征学习,检测可被文字描述的PII线索,提升行人重识别(Re-ID)性能。该方法聚焦可解释特征,超越低层外观线索,实现语义层面的隐私识别。我们系统评估了多个行人图像数据集中PII的存在情况,实验表明在跨数据集场景下(如Market-1501到CUHK03-np)性能显著提升,验证了框架的实际价值。代码已开源。

原文摘要 · Abstract (English)

The collection and release of street-level recordings as Open Data play a vital role in advancing autonomous driving systems and AI research. However, these datasets pose significant privacy risks, particularly for pedestrians, due to the presence of Personally Identifiable Information (PII) that extends beyond biometric traits such as faces. In this paper, we present cRID, a novel cross-modal framework combining Large Vision-Language Models, Graph Attention Networks, and representation learning to detect textual describable clues of PII and enhance person re-identification (Re-ID). Our approach focuses on identifying and leveraging interpretable features, enabling the detection of semantically meaningful PII beyond low-level appearance cues. We conduct a systematic evaluation of PII presence in person image datasets. Our experiments show improved performance in practical cross-dataset Re-ID scenarios, notably from Market-1501 to CUHK03-np (detected), highlighting the framework's practical utility. Code is available at https://github.com/RAufschlaeger/cRID.

行人重识别隐私保护跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。