让AI看图识人时能说出依据,提升可解释性。
InterPartAbility: Phrase-Region Grounding for Interpretable Text-to-Image Person Re-Identification

- 用概念级短语引导模型关注图像局部区域,实现精准定位。
- 在三个基准上同时达到顶尖检索准确率与可解释性表现。
- 适合需要透明决策过程的安防、医疗等高风险场景使用。
文本到图像的人体重识别(TI-ReID)依赖自然语言描述从参考图像库中检索匹配个体。尽管大型视觉语言模型(VLMs)已取得优异检索性能,但其决策过程缺乏可解释性。现有可解释性方法仅依赖槽注意力突出关注区域,无法可靠关联视觉区域与语义概念,且解释范围受限于有限词汇。本文提出InterPartAbility,一种可解释的TI-ReID方法,实现显式的部件级匹配与短语-区域对齐。不同于参数量大的槽注意力方法,本方法通过开放词汇的补丁-短语交互模块(PPIM)以概念级短语指导标准TI-ReID模型,使模型更关注对应局部区域。进一步利用CLIP ViT自注意力生成空间集中化的补丁激活,形成与各部件短语对齐的解释地图。此外,提出适用于TI-ReID的定量可解释性评估协议,引入反事实区域移除实验,衡量当最高评分解释区域被移除后的检索性能下降。在三个挑战性基准上的实证结果表明,InterPartAbility在该评估指标下实现最先进可解释性表现,同时保持具有竞争力的检索准确性。
原文摘要 · Abstract (English)
Text-to-image person re-identification (TI-ReID) relies on natural-language text descriptions to retrieve top matching individuals from a gallery of reference images. While recent large vision-language models (VLMs) achieve strong retrieval performance, their decisions remain largely uninterpretable. Existing interpretability approaches in TI-ReID rely solely on slot-attention to highlight attended regions, but fail to reliably bind visual regions to semantically meaningful concepts, limiting interpretation to qualitative visualizations over a restricted vocabulary. This paper introduces InterPartAbility, an interpretable TI-ReID method that performs explicit part-wise matching and enables phrase-region grounding. Unlike parameter-heavy slot-attention methods that yield only qualitative interpretability, our open-vocabulary patch-phrase interaction module (PPIM) guides a standard TI-ReID model with concept-level phrases. Concept-based part phrases provide evidence that encourages the model to attend to the corresponding local image regions. InterPartAbility further leverages CLIP ViT self-attention to produce spatially concentrated patch activations aligned with each part-level phrase, yielding grounded explanation maps. Finally, a quantitative interpretability protocol for TI-ReID is introduced that extends current perturbation-based evaluation metrics into the TI-Reid domain. This includes a counterfactual region removal that measures retrieval degradation when top-ranked explanatory regions are removed. Empirical results on three challenging benchmarks show that InterPartAbility can achieve SOTA interpretability performance under these metrics, while sustaining competitive retrieval accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。