用场景信息提升文本搜人的准确率,解决描述模糊和环境复杂问题
SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking
- 分两阶段检索:先对齐文本与行人区域,再用场景感知重排
- 在13万张真实场景图上验证,性能显著优于现有方法
- 无需训练的重排模块,适合实际应用中快速优化结果
文本驱动的人体检索旨在通过自然语言描述从图像画廊中识别目标个体。现有方法主要依赖外观特征进行跨模态匹配,但受限于场景视觉复杂性和文本描述的固有模糊性。上下文信息(如地标、关系线索)可提供补充线索,却未被充分挖掘。为此,我们提出一种新范式:场景感知的文本驱动人体检索,显式融合个体外观与全局场景上下文以提升检索精度。首先构建了包含超10万张真实场景、涵盖行人属性与场景上下文的大型基准数据集ScenePerson-13W。基于此,提出SA-Person两阶段框架:第一阶段通过外观定位对齐文本与行人群体区域;第二阶段引入无需训练的SceneRanker模块,联合推理行人外观与全局场景上下文进行重排。在ScenePerson-13W及现有基准上的大量实验验证了其有效性。数据集与代码将公开发布,推动后续研究。
原文摘要 · Abstract (English)
Text-based person retrieval aims to identify a target individual from an image gallery using a natural language description. Existing methods primarily focus on appearance-driven cross-modal retrieval, yet face significant challenges due to the visual complexity of scenes and the inherent ambiguity of textual descriptions. The contextual information, such as landmarks and relational cues, provides complementary cues that can offer valuable complementary insights for retrieval, but remains underexploited in current approaches. Motivated by this limitation, we propose a novel paradigm: scene-aware text-based person retrieval, which explicitly integrates both individual appearance and global scene context to improve retrieval accuracy. To support this, we first introduce ScenePerson-13W, a large-scale benchmark dataset comprising over 100,000 real-world scenes with rich annotations encompassing both pedestrian attributes and scene context. Based on this dataset, we further present SA-Person, a two-stage retrieval framework. In the first stage, SA-Person performs discriminative appearance grounding by aligning textual descriptions with pedestrian-specific regions. In the second stage, it introduces SceneRanker, a training-free, scene-aware re-ranking module that refines retrieval results by jointly reasoning over pedestrian appearance and the global scene context. Extensive experiments on ScenePerson-13W and existing benchmarks demonstrate the effectiveness of our proposed SA-Person. Both the dataset and code will be publicly released to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。