解决笼中动物追踪与姿态估计的遮挡难题
UnCageNet: Tracking and Pose Estimation of Caged Animal
- 用增强方向感知的网络分割笼结构,识别遮挡区域
- 通过内容感知修复填补笼子遮挡部分,恢复完整图像
- 显著提升关键点检测精度和轨迹一致性,适合动物行为研究
动物追踪与姿态估计系统(如STEP和ViTPose)在存在笼子结构和系统性遮挡的图像与视频中性能大幅下降。本文提出三阶段预处理流程:(1) 使用带可调方向滤波器的Gabor增强ResNet-UNet进行笼子分割;(2) 采用CRFill实现内容感知的遮挡区域重建;(3) 在去笼图像上评估姿态估计与追踪效果。所提Gabor增强分割模型利用72个方向核提取方向敏感特征,精准识别严重干扰现有方法的笼结构。实验表明,通过本流程去除笼子遮挡后,姿态估计与追踪性能接近无遮挡环境表现,关键点检测准确率与轨迹一致性均显著提升。
原文摘要 · Abstract (English)
Animal tracking and pose estimation systems, such as STEP (Simultaneous Tracking and Pose Estimation) and ViTPose, experience substantial performance drops when processing images and videos with cage structures and systematic occlusions. We present a three-stage preprocessing pipeline that addresses this limitation through: (1) cage segmentation using a Gabor-enhanced ResNet-UNet architecture with tunable orientation filters, (2) cage inpainting using CRFill for content-aware reconstruction of occluded regions, and (3) evaluation of pose estimation and tracking on the uncaged frames. Our Gabor-enhanced segmentation model leverages orientation-aware features with 72 directional kernels to accurately identify and segment cage structures that severely impair the performance of existing methods. Experimental validation demonstrates that removing cage occlusions through our pipeline enables pose estimation and tracking performance comparable to that in environments without occlusions. We also observe significant improvements in keypoint detection accuracy and trajectory consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。