arXiv:2509.09118cs.CV2025-09EMNLP被引 6

用新数据集和双掩码框架提升文本行人检索的准确性

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

  • 用大模型自动构建500万张行人图文对数据集
  • 通过梯度注意力动态屏蔽噪声文本,提升跨模态对齐
  • 适合做行人重识别与多模态学习的研究者

尽管对比语言图像预训练(CLIP)在多种视觉任务中表现优异,但在行人表征学习中面临两大挑战:一是缺乏大规模聚焦行人的视觉-语言标注数据;二是全局对比学习难以保持细粒度匹配所需的局部特征,且易受噪声文本标记影响。本文通过数据构建与模型架构协同优化推进CLIP在行人表征学习中的应用。首先,利用多模态大模型的上下文学习能力,开发了抗噪声的数据构建流程,自动筛选并标注网络获取的图像,生成包含500万条高质量行人图文对的WebPerson数据集。其次,提出GA-DMS(梯度注意力引导双掩码协同框架),基于梯度注意力相似度自适应屏蔽噪声文本标记,增强跨模态对齐;同时引入掩码标记预测目标,促使模型学习更具信息量的文本语义表示。大量实验表明,GA-DMS在多个基准上达到当前最优性能。

原文摘要 · Abstract (English)

Although Contrastive Language-Image Pre-training (CLIP) exhibits strong performance across diverse vision tasks, its application to person representation learning faces two critical challenges: (i) the scarcity of large-scale annotated vision-language data focused on person-centric images, and (ii) the inherent limitations of global contrastive learning, which struggles to maintain discriminative local features crucial for fine-grained matching while remaining vulnerable to noisy text tokens. This work advances CLIP for person representation learning through synergistic improvements in data curation and model architecture. First, we develop a noise-resistant data construction pipeline that leverages the in-context learning capabilities of MLLMs to automatically filter and caption web-sourced images. This yields WebPerson, a large-scale dataset of 5M high-quality person-centric image-text pairs. Second, we introduce the GA-DMS (Gradient-Attention Guided Dual-Masking Synergetic) framework, which improves cross-modal alignment by adaptively masking noisy textual tokens based on the gradient-attention similarity score. Additionally, we incorporate masked token prediction objectives that compel the model to predict informative text tokens, enhancing fine-grained semantic representation learning. Extensive experiments show that GA-DMS achieves state-of-the-art performance across multiple benchmarks.

行人检索多模态CLIP数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。