用注意力机制自动发现并对齐图像与文本中的人体部位,提升跨模态检索精度。
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery

- 基于槽注意力的部件发现模块,无需人工标注即可自动识别关键身体部位
- 在三个公开数据集上显著超越现有方法,最高提升超过10%的检索准确率
- 适合需要高可解释性与精准定位的跨模态人物搜索场景
基于文本的人物搜索通过自由文本查询在大规模图像集合中定位特定个体,面临视觉与文本表征在人体部位层面难以对齐的挑战。现有方法因缺乏直接的部位级监督且依赖启发式特征,在部件特征提取与对齐方面表现不佳。本文提出一种新框架,利用基于槽注意力的部件发现模块,自主识别并跨模态对齐显著性身体部位,提升可解释性与检索准确率,且无需显式的部位对应标注。此外,文本驱动的动态部件注意力机制进一步调整各部位重要性,优化检索效果。该方法在三个公开基准上进行评估,显著优于现有方法。
原文摘要 · Abstract (English)
Text-based person search, employing free-form text queries to identify individuals within a vast image collection, presents a unique challenge in aligning visual and textual representations, particularly at the human part level. Existing methods often struggle with part feature extraction and alignment due to the lack of direct part-level supervision and reliance on heuristic features. We propose a novel framework that leverages a part discovery module based on slot attention to autonomously identify and align distinctive parts across modalities, enhancing interpretability and retrieval accuracy without explicit part-level correspondence supervision. Additionally, text-based dynamic part attention adjusts the importance of each part, further improving retrieval outcomes. Our method is evaluated on three public benchmarks, significantly outperforming existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。