用连续元数据增强模型,让动物识别更抗时间变化。
Parameter-Efficient Vision-Language Adaptation with Continuous Metadata Conditioning for Animal Re-Identification

- 用连续数值元数据直接调节视觉表征,不需离散化。
- 在七年鱼类数据集上准确率提升12.3%,跨场景表现更好。
- 参数高效,推理时无需元数据,适合野外长期监测。
长期动物重识别需应对形态演变和季节性外观变化。尽管视觉语言模型具备强预训练视觉表示,但在长期生态场景下的适应仍具挑战,尤其面对身份与时间分布漂移。本文提出一种参数高效的CLIP适配框架,引入连续元数据条件机制,在训练中将数值属性直接融入提示表征。结合低秩视觉适配、基于提示的监督与跨模态对齐,该方法通过保留数值元数据的连续结构而非离散化为文本类别,实现嵌入空间的平滑调制,同时保持纯视觉推理流程。在七年的鱼类纵向数据集及多个野生动物基准测试上,闭集、开集与时间感知评估均取得改进。结果表明,连续元数据条件显著提升对长期外观变化与时间分布漂移的鲁棒性,参数高效适配支持测试时无需元数据的纯视觉推理。代码与评估划分见:https://github.com/AnilOsmanTur/MetaPrompt-ReID。
原文摘要 · Abstract (English)
Long-term animal re-identification (ReID) must remain robust to gradual morphological evolution and seasonal appearance shifts. Although recent vision-language models provide strong pretrained visual representations, adapting them to longitudinal ecological settings remains challenging, particularly under identity and temporal distribution shifts. We present a parameter-efficient CLIP adaptation framework for animal ReID and introduce a continuous metadata-conditioning mechanism that incorporates numerical attributes directly into the prompt representation during training. While low-rank visual adaptation, prompt-based supervision, and cross-modal alignment provide the adaptation framework, the proposed metadata-conditioning strategy constitutes the primary methodological contribution. By preserving the continuous structure of numerical metadata rather than discretizing it into textual categories, the proposed approach enables smooth modulation of the embedding space during training while maintaining a purely visual inference pipeline. Experiments on a seven-year longitudinal fish dataset and multiple wildlife benchmarks demonstrate improved performance under closed-set, open-set, and time-aware evaluation protocols. The results demonstrate that continuous metadata conditioning improves robustness to longitudinal appearance variation and temporal distribution shifts, while parameter-efficient adaptation enables a purely visual inference pipeline without requiring metadata at test time. Code and evaluation splits can be found at: https://github.com/AnilOsmanTur/MetaPrompt-ReID.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。