解决低分辨率监控中文本检索图像的可靠性与排序漂移问题
Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance

- 通过分辨率为条件的推理模块评估视觉特征可靠性
- 在超低分辨率下实现平均5.7%的Rank-1提升
- 适合实际监控场景下的跨分辨率文本图像检索
文本到图像人员识别(TIPR)利用自然语言描述检索目标人物。然而,现有方法大多忽视真实监控场景中的分辨率差异。本文将跨分辨率TIPR的问题归结为两个耦合的失效模式:证据可靠性崩溃(ERC),即退化的视觉标记难以支撑细粒度文本定位;以及排名分布漂移(RDD),即混合分辨率图库扭曲相似性邻域,导致检索排名不稳定。为此,我们提出跨分辨率语义迁移(CRST),一个类CLIP框架,包含三个模块:分辨率为条件的推理、文本引导的精炼和CR-RDA。分辨率为条件的推理模块估计标记可靠性,抑制受损证据;文本引导的精炼模块注入语义先验,恢复区分性线索;CR-RDA模块将高分辨率邻域几何结构迁移至低分辨率环境,稳定混合分辨率下的检索排名。在CUHK-PEDES、ICFG-PEDES和RSTPReid数据集上的实验表明,CRST在超低分辨率下平均提升Rank-1达5.7%,mAP提升5.3%,且在不牺牲高分辨率性能的前提下稳定了混合分辨率检索。代码将公开。
原文摘要 · Abstract (English)
Text-to-image person re-identification (TIPR) retrieves target persons using natural language descriptions. However, existing methods largely overlook resolution variance in real-world surveillance. They characterize cross-resolution TIPR through two coupled failure modes: Evidence Reliability Collapse (ERC), where degraded visual tokens become unreliable for grounding fine-grained text, and Ranking Distribution Drift (RDD), where mixed-resolution galleries distort similarity neighborhoods and destabilize retrieval rankings. To address this challenge, we propose Cross-Resolution Semantic Transfer (CRST), a CLIP-style framework with three modules: resolution-conditioned reasoning, text-guided refinement and CR-RDA. Resolution-conditioned reasoning estimates token reliability to suppress corrupted evidence. Text-guided refinement injects semantic priors to recover discriminative cues. CR-RDA transfers HR neighborhood geometry to stabilize LR ranking under mixed resolutions. Experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReid show that CRST improves ultra-low-resolution Rank-1 and mAP on average by 5.7% and 5.3%, while stabilizing mixed-resolution retrieval without sacrificing high-resolution accuracy.The code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。