ViT模型能自动学习动物性别年龄等生物特征,且可解释。
What Does Animal Re-Identification Learn? Linear Biological Concepts and Their Origins in Visual Representations

- 用对比损失微调ViT,发现性别年龄在特征空间中呈线性分布。
- 单图即可识别性别,最高达0.91 AUROC,且可被激活操控翻转预测。
- 模型不创造新概念,而是重新定位已有生物特征,适合生态监控可解释性研究。
保护工作日益依赖相机陷阱收集海量野生动物图像,远超人工分析能力,因此个体动物重识别(Re-ID)对种群监测至关重要。然而,基于视觉变压器(ViT)的Re-ID模型缺乏对生物概念的显式监督,难以理解其决策依据。我们探究这类模型是否仍能在表征中组织出具有生物学意义的轴线。以DINOv3为骨干网络,使用三元组间隔损失微调西低地大猩猩Re-ID任务,发现性别与年龄在特征空间中呈现线性方向,且在未见个体上表现良好,单张图像即可实现最高0.91 AUROC的性别分类。激活操纵实验表明,性别方向对模型决策具有因果作用,可显著改变预测结果。对比预训练与微调骨干网络发现,Re-ID训练并未生成这些概念,而是将其重新映射至网络不同位置。数据归因分析显示,所提取的表征反映连续的生物学维度,具有冗余编码,并受视觉模糊个体影响。这些发现揭示了Re-ID表征中的生物结构与失效模式,推动野生动物监控中可审计计算机视觉的发展。
原文摘要 · Abstract (English)
Conservation increasingly relies on camera traps that collect more wildlife imagery than experts can manually analyze, making animal re-identification (Re-ID) essential for monitoring individuals and populations. Yet understanding which cues drive model decisions is challenging for ViT-based Re-ID models, whose metric-learning objectives provide no explicit supervision for biological concepts. We ask whether such models nonetheless organize their representations along biologically meaningful axes. Using a DINOv3 backbone fine-tuned for Western lowland gorilla Re-ID with triplet-margin loss, we find that sex and age emerge as linear directions that generalize to held-out individuals, reaching up to 0.91 AUROC and being recoverable from a single image per individual. Activation steering further shows that the sex direction is causally used by the model, flipping a significant fraction of predictions to the opposite sex. Comparing off-the-shelf and fine-tuned backbones shows that Re-ID training does not create these concepts, but relocates them across the network. Finally, data attribution reveals that the representation we find reflects a graded biological axis, is redundantly encoded across the population and shaped by visually ambiguous individuals. Together, these findings show how interpretability can uncover both the biological structure and failure modes of Re-ID representations, providing a step toward auditable computer vision for wildlife monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。