arXiv:2601.14978cs.CV2026-01

一个模型搞定多个数据集的文本行人检索,解决跨数据集适应难题。

Unified Multi-Dataset Training for TBPS

  • 统一清洗多源数据,自动过滤噪声图文对
  • 单模型在5个数据集上超越独立训练模型
  • 适合需要跨数据集部署的行人检索应用

文本驱动的行人检索(TBPS)虽借助视觉语言模型取得进展,但受限于训练数据不足,且现有模型未针对行人识别预训练。当前方法需为每个数据集单独微调,导致多个独立模型。尽管合成数据可扩充规模,仍无法消除数据集特异性。本文提出Scale-TBPS,通过两个创新:(i) 噪声感知的统一数据集整合策略;(ii) 可扩展的判别性身份学习框架,有效应对大量唯一身份。在CUHK-PEDES、ICFG-PEDES、RSTPReid、IIITD-20K和UFine6926共5个数据集上的实验表明,单一Scale-TBPS模型性能优于各数据集优化模型及简单联合训练。

原文摘要 · Abstract (English)

Text-Based Person Search (TBPS) has seen significant progress with vision-language models (VLMs), yet it remains constrained by limited training data and the fact that VLMs are not inherently pre-trained for pedestrian-centric recognition. Existing TBPS methods therefore rely on dataset-centric fine-tuning to handle distribution shift, resulting in multiple independently trained models for different datasets. While synthetic data can increase the scale needed to fine-tune VLMs, it does not eliminate dataset-specific adaptation. This motivates a fundamental question: can we train a single unified TBPS model across multiple datasets? We show that naive joint training over all datasets remains sub-optimal because current training paradigms do not scale to a large number of unique person identities and are vulnerable to noisy image-text pairs. To address these challenges, we propose Scale-TBPS with two contributions: (i) a noise-aware unified dataset curation strategy that cohesively merges diverse TBPS datasets; and (ii) a scalable discriminative identity learning framework that remains effective under a large number of unique identities. Extensive experiments on CUHK-PEDES, ICFG-PEDES, RSTPReid, IIITD-20K, and UFine6926 demonstrate that a single Scale-TBPS model outperforms dataset-centric optimized models and naive joint training.

行人检索多数据集视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。