用自动标注提升文本搜人模型的细粒度对齐能力
ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search

- 通过自动匹配区域与句子生成伪标注,减少人工标注依赖
- 融合全局对比与局部对齐,在长句查询上提升显著
- 新基准P-VLG含超10万区域标注,支持多粒度评估
文本驱动行人搜索(TBPS)旨在通过自然语言查询检索行人图像。现有基于CLIP的模型因全局表征偏见和短标题训练带来的语义稀疏性,难以实现细粒度理解,导致对齐能力弱,且区域级标注稀缺加剧此问题。为此,本文提出ROGLE(鲁棒全局-局部嵌入)框架,通过自动区域-句子匹配(RSM)策略自动生成伪区域-句子对,实现可扩展的细粒度监督。ROGLE采用多粒度学习机制,融合全局对比学习与区域级局部对齐。同时构建了大尺度的P-VLG基准,从公开数据集整理并扩充图像,包含超过10万标注区域和丰富长句描述,是首个支持全局与局部评估协议的TBPS基准。大量实验表明,ROGLE显著优于现有方法,尤其在长句查询下表现突出。代码与数据集将公开。
原文摘要 · Abstract (English)
Text-Based Person Search (TBPS) aims to retrieve pedestrian images using natural language queries. However, existing TBPS models, especially those based on CLIP, struggle with fine-grained understanding due to global representational bias and semantic sparsity inherited from training on short captions. This results in weak fine-grained alignment, exacerbated by the scarcity of region-level annotations. To address this, we propose ROGLE (Robust Global-Local Embedding), a unified framework that overcomes reliance on costly manual annotations through an automated Region-to-Sentence Matching (RSM) strategy. RSM automatically mines pseudo region-sentence pairs for scalable fine-grained supervision. Furthermore, ROGLE employs a multi-granular learning strategy that fuses global contrastive learning with region-level local alignment. We also introduce the P-VLG Benchmark, a large-scale dataset constructed by curating and enriching images from established public benchmarks. It features over 100,000 annotated regions and rich long-form captions, making it the first TBPS benchmark to support both global and local assessment protocols. Extensive experiments show that ROGLE significantly outperforms existing approaches, particularly on challenging long-form queries. Code and the P-VLG benchmark will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。