arXiv:2603.02270cs.CV2026-03中稿 · CVPR

用合成文本提升宠物识别准确率,效果比纯视觉方法高11%。

From Visual to Multimodal: Systematic Ablation of Encoders and Fusion Strategies in Animal Identification

  • 融合视觉与合成文本描述,用门控机制整合多模态特征。
  • 在190万张照片上测试,达到84.28%准确率和0.0422的错误率。
  • 适合做大规模宠物识别、跨模态学习研究者参考。

自动宠物识别对找回走失宠物有实际价值,但现有系统受限于数据规模小且依赖单一视觉线索。本研究提出一种多模态验证框架,通过合成文本描述引入语义身份先验来增强视觉特征。构建了包含190万张照片、覆盖695,091只唯一动物的大规模训练数据集。系统性消融实验表明,SigLIP2-Giant和E5-Small-v2分别为最优视觉与文本骨干网络。评估了从简单拼接到自适应门控等多种融合策略,最终采用门控融合机制,在综合测试协议下实现84.28%的Top-1准确率和0.0422的等错误率。相比领先单模态基线提升11%,证明合成语义描述能显著优化大规模宠物重识别中的决策边界。

原文摘要 · Abstract (English)

Automated animal identification is a practical task for reuniting lost pets with their owners, yet current systems often struggle due to limited dataset scale and reliance on unimodal visual cues. This study introduces a multimodal verification framework that enhances visual features with semantic identity priors derived from synthetic textual descriptions. We constructed a massive training corpus of 1.9 million photographs covering 695,091~unique animals to support this investigation. Through systematic ablation studies, we identified SigLIP2-Giant and E5-Small-v2 as the optimal vision and text backbones. We further evaluated fusion strategies ranging from simple concatenation to adaptive gating to determine the best method for integrating these modalities. Our proposed approach utilizes a gated fusion mechanism and achieved a Top-1 accuracy of 84.28\% and an Equal Error Rate of 0.0422 on a comprehensive test protocol. These results represent an 11\% improvement over leading unimodal baselines and demonstrate that integrating synthesized semantic descriptions significantly refines decision boundaries in large-scale pet re-identification.\\\\ \textbf{Conference update.} This work was accepted to the FGVC13 Workshop at CVPR 2026. We thank Vadim Ezhov and Nikolai Shcheglov for their help with the poster presentation and demonstration at the conference.

宠物识别多模态门控融合图像检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。