首个面向非洲草原野生动物的无人机单目3D检测数据集与基准测试
WildBox: A Dataset and Benchmark for Aerial Monocular 3D Detection of African Savanna Wildlife

- 构建了含23.7万标注框的多物种无人机视频数据集,支持跨片段实例追踪
- 零样本3D检测性能归零,但微调后达到13.17的平均AP3D,深度估计是主要瓶颈
- 提出粗到精训练策略,提升斑马亚种识别效果,适合野生动物监测研究者
我们推出WildBox,一个用于无人机视频中野生动物单目3D检测的数据集与基准测试,包含七个非洲草原物种、六个基准类别,共237,505个3D边界框标注。标注采用与KITTI/Omni3D兼容的格式,按视频段落进行尺度归一化,保持实例身份连续性。评估两种开放词汇单目3D架构:OVMono3D-LIFT与DetAny3D,涵盖零样本、真实2D框提示及监督微调三种协议。开放词汇2D基础模型可实现50.55 AP@50的零样本定位,但零样本3D检测在所有条件下(包括真实2D框输入)均退化至0.00 AP,表明失败集中于3D阶段。在WildBox上微调后,性能恢复至8.68±0.47 [email protected]和13.17±0.69 AP3D宏平均。微调后深度贡献了84%的归一化豪斯多夫距离,零样本下超过99%,凸显单目航拍深度估计为该任务的核心挑战。采用粗到精课程学习(先合并斑马类预训练,再细分为Grevy's/平原斑马微调),在降低总计算量的同时显著提升宏平均3D性能,尤其在两个斑马子类上增益最大。WildBox已发布视频级划分、评测代码及基线检查点,推动无人机视频中3D野生动物感知的发展。
原文摘要 · Abstract (English)
We introduce WildBox, a dataset and benchmark for monocular 3D detection of wildlife from drone video, comprising 237,505 3D bounding box annotations across seven African savanna species grouped into six benchmark classes. Annotations follow a KITTI/Omni3D-compatible format in a per-segment scale-normalised camera frame, with instance identities maintained across each segment. We evaluate two open-vocabulary monocular 3D architectures, OVMono3D-LIFT and DetAny3D, under zero-shot, ground-truth 2D box prompt, and supervised fine-tuning protocols. Open-vocabulary 2D foundation models provide usable zero-shot wildlife localisation (50.55 AP@50), but zero-shot 3D detection collapses to 0.00 AP across both architectures and every 2D-input condition tested, including ground-truth 2D box prompts, thus isolating the failure to the 3D stage. Fine-tuning on WildBox recovers performance to 8.68 +/- 0.47 [email protected] and 13.17 +/- 0.69 AP3D macro. Depth contributes 84% of normalised Hausdorff distance after fine-tuning and over 99% in zero-shot, identifying monocular aerial depth as the dominant open problem in this regime. A coarse-to-fine curriculum, i.e. pretraining on a merged zebra class before fine-tuning on the Grevy's/plains split, improves macro 3D performance with less total compute, with the largest gains on the two zebra subclasses. WildBox is released with video-level splits, evaluation code, and baseline checkpoints to enable progress in 3D wildlife perception from drone video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。