arXiv:2507.10195cs.CV2025-07被引 8

通过双层次对齐缩小图文人物检索域差距,提升真实场景效果

Minimizing the Pretraining Gap: Domain-aligned Text-Based Person Retrieval

  • 图像与区域级双通道域适应,对齐合成与真实数据分布
  • 在CUHK-PEDES等3个数据集上达到最新最优性能
  • 适合关注跨域图文检索、真实场景应用的研究者

本文聚焦于基于文本的人物检索任务,即根据文本描述识别个体。尽管合成数据预训练取得进展,但光照、色彩、视角等差异导致的域差距仍限制了预训练-微调范式的有效性。为此,我们提出统一管道,在图像和区域两个层面实现域自适应。方法包含两项核心组件:图像级的领域感知扩散(DaD),用于对齐合成与真实世界的图像分布;区域级的多粒度关系对齐(MRA),用于对齐视觉区域与文本描述,解决细粒度差异。该双层级策略有效弥合域差距,在CUHK-PEDES、ICFG-PEDES和RSTPReid三个数据集上实现当前最佳表现。代码与模型已开源。

原文摘要 · Abstract (English)

In this work, we focus on text-based person retrieval, which identifies individuals based on textual descriptions. Despite advancements enabled by synthetic data for pretraining, a significant domain gap, due to variations in lighting, color, and viewpoint, limits the effectiveness of the pretrain-finetune paradigm. To overcome this issue, we propose a unified pipeline incorporating domain adaptation at both image and region levels. Our method features two key components: Domain-aware Diffusion (DaD) for image-level adaptation, which aligns image distributions between synthetic and real-world domains, e.g., CUHK-PEDES, and Multi-granularity Relation Alignment (MRA) for region-level adaptation, which aligns visual regions with descriptive sentences, thereby addressing disparities at a finer granularity. This dual-level strategy effectively bridges the domain gap, achieving state-of-the-art performance on CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets. The dataset, model, and code are available at https://github.com/Shuyu-XJTU/MRA.

图文检索域适应人物识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。