通过局部对齐提升鱼类重识别准确率
SLAP: Selective Local Vision-Language Alignment for Fish Re-Identification via Partial Optimal Transport

- 用部分最优传输实现视觉块与身份提示的选区匹配
- 在多组鱼类数据集上显著优于现有CLIP方法
- 适合细粒度生物识别与海洋动物追踪场景
个体鱼类重识别(ReID)是一项细粒度识别任务,其身份判别线索通常集中于特定身体区域而非均匀分布。然而,现有基于CLIP的ReID方法主要依赖全局图像-文本对齐,导致背景和弱判别区域也参与跨模态监督。本文提出一种选择性局部视觉-语言对齐框架,通过部分最优传输(POT)建立视觉块嵌入与多个身份感知提示嵌入之间的局部对应关系。POT不强制所有对应,而是允许模型聚焦最强跨模态匹配,避免弱匹配区域的强行对齐,从而生成更具判别性的视觉表示用于检索。该框架端到端训练,推理时仅保留适配后的视觉编码器。在纵向的Symphodus melops数据集上,无论闭集还是开集评估协议下均持续优于近期基于CLIP的ReID方法。在其他数据集上的额外实验进一步验证了该方法在多种海洋ReID基准上的泛化能力。
原文摘要 · Abstract (English)
Individual fish re-identification (ReID) is a fine-grained recognition problem in which identity-discriminative cues are often localized to specific body regions rather than distributed uniformly across the animal. Nevertheless, recent CLIP-based ReID methods rely predominantly on global image-text alignment, allowing background and weakly discriminative regions to contribute to cross-modal supervision. We propose a selective local vision-language alignment framework that establishes localized correspondences between visual patch embeddings and multiple identity-aware prompt embeddings through Partial Optimal Transport (POT). Rather than enforcing exhaustive correspondence, POT enables selective matching between visual patches and prompt embeddings, allowing the model to emphasize the strongest cross-modal correspondences while avoiding forced alignment of weakly matching regions, thereby yielding more discriminative visual representations for retrieval. The framework is trained end-to-end, while only the adapted visual encoder is retained during inference. Experiments on the longitudinal Symphodus melops dataset demonstrate consistent improvements over recent CLIP-based ReID methods under both closed-set and open-set evaluation protocols. Additional evaluations on other datasets further demonstrate the generalization capability of the proposed method across diverse marine ReID benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。