解决文本检索中查询粒度不确定的问题,实现任意粒度的精准人物检索。
Achieving Text-based Person Retrieval with Any Granularity

- 提出五级粒度谱系与多粒度标注工具,构建高质量数据集
- 设计新评估框架MG-Eval,支持从粗到细的渐进式查询评测
- 创新跨模态多粒度对齐框架,提升不同粒度下的检索性能
文本驱动的人物检索面临一个关键但未被充分探索的挑战:现实场景中查询粒度存在固有不确定性。本文提出‘任意粒度文本人物检索’新范式,并提供系统性解决方案。首先,定义五级粒度谱系,利用新型多粒度文本标注引擎构建高质多粒度数据集UFine6926-MG。其次,鉴于粗粒度查询天然对应多个有效候选者,提出MG-Eval评估基准,包含渐进式细化文本与跨身份标签,反映真实语义,并配套定制化评估指标与协议。第三,通过全面诊断发现现有方法系统性局限后,提出跨模态多粒度对齐与匹配(CMAM)框架。该框架通过:1)正交专家感知分离粒度特定特征;2)概率对齐建模查询不确定性下的多对多匹配;3)粒度一致性推理,通过跨模态联合验证引导特征学习。实验表明,CMAM在所有粒度层级上显著优于现有最先进方法。本工作建立了基础基准与稳健基线,为更实用的人物检索系统铺平道路。
原文摘要 · Abstract (English)
Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This paper introduces a new paradigm, Text-based Person Retrieval with Any Granularity, and provides a systematic solution. First, we formalize a five-level granularity spectrum and construct UFine6926-MG, a high-quality multi-grained dataset annotated comprehensively at all granularities via a novel Multi-grained Text Annotation Engine. Second, acknowledging that coarse queries naturally correspond to multiple valid candidates, we propose MG-Eval, a holistic evaluation benchmark with progressively detailed texts and cross-identity labels that reflect real-world semantics, alongside tailored evaluation metrics and protocols. Third, after a comprehensive diagnosis reveals the systemic limitations of existing research, we propose the Cross-modal Multi-grained Aligning and Matching (CMAM) framework. CMAM achieves granularity-aware retrieval through: 1) orthogonal-expert perception to disentangle granularity-specific features; 2) probabilistic alignment to model many-to-many matches under query uncertainty; and 3) granularity-consistent reasoning to steer feature learning via joint cross-modal granularity verification. Experiments demonstrate that CMAM significantly outperforms state-of-the-art methods across all granularity levels. This work establishes a foundational benchmark and a robust baseline, paving the way for more practical person retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。