融合外观与运动特征,提升企鹅检测与识别准确率。
Detection and Identification of Penguins Using Appearance and Motion Features
- 用双帧输入改进YOLO11,利用运动信息增强检测
- [email protected]从0.922提升至0.933,可识别静态图像中无法区分的个体
- 基于轨迹的对比学习减少身份切换,适合动物行为追踪场景
在动物园区中,对企鹅进行持续监控至关重要但技术难度高,原因在于其视觉特征高度相似、姿态变化频繁且环境噪声大(如水波反射)。本文提出一种融合外观与运动特征的框架,以提升检测与识别性能。针对检测任务,将YOLO11改进为处理连续帧,利用运动线索在视觉特征模糊时仍能定位目标。实验表明,使用双帧输入微调模型后,[email protected]从0.922提升至0.933,显著优于基线,并成功恢复静态图像中无法区分的个体。针对识别任务,提出一种跟踪轨迹后应用的对比学习方法。定性可视化显示,该方法生成一致的特征嵌入,同一个体样本在特征空间中更靠近,表明有潜力缓解身份切换问题。
原文摘要 · Abstract (English)
In animal facilities, continuous surveillance of penguins is essential yet technically challenging due to their homogeneous visual characteristics, rapid and frequent posture changes, and substantial environmental noise such as water reflections. In this study, we propose a framework that enhances both detection and identification performance by integrating appearance and motion features. For detection, we adapted YOLO11 to process consecutive frames to overcome the lack of temporal consistency in single-frame detectors. This approach leverages motion cues to detect targets even when distinct visual features are obscured. Our evaluation shows that fine-tuning the model with two-frame inputs improves [email protected] from 0.922 to 0.933, outperforming the baseline, and successfully recovers individuals that are indistinguishable in static images. For identification, we introduce a tracklet-based contrastive learning approach applied after tracking. Through qualitative visualization, we demonstrate that the method produces coherent feature embeddings, bringing samples from the same individual closer in the feature space, suggesting the potential for mitigating ID switching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。